Release: development → main (442 commits) — G5 of the runbook, HOLD at G4b - #410
Polichinel wants to merge 497 commits into
Conversation
…tion docs(liveness): naming convention — assumption upgraded to cited fact (views_api wiki)
…ndows (#240) python -m tools.liveness.datafactory_input answers: does the input store's observed coverage (live last_valid_month_id) reach what meta/partitions.json requires? Requirement DERIVED at runtime (max test-window end — re-arms on every partition bump automatically; automates the C-96 tripwire). TDD (14 tests first, injected reader/netrc seams); netrc presence reported as a fact, values never read; exit 0/1/2. First live run: INPUT_FRESH (558 vs 552, margin 6 months). Also: S1 review-finding remediation — parse_run_name now rejects legacy month-00 names (fatalities001_2022_00_t01) with a pinned test, keeping the month math clean. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…tory-input feat(liveness): S2 — datafactory input check (coverage vs partition windows) (#240)
…ion IDs encoded (#241) python -m tools.liveness.appwrite_store answers: is the new shelf reachable, when did the newest forecast land, and do the REAL metadata IDs hold? The IDs are now discovered + encoded with receipts (db 'file_metadata', collection 'production_forecasts'); the June 2026 failure's phantom 'forecasts_metadata' is documented as the historical wrong value — closing C-100's founding mystery. TDD (19 tests first): injected fetch/credentials/ clock; creds resolved env-first then known .env (ancestor-walk discovery — robust from worktrees); secrets never rendered (pinned redaction test); truthful SKIP without creds. Verdicts STORE_ACTIVE/IDLE/UNREACHABLE/SKIP, exit 0/1/2. First live run: STORE_IDLE — newest file 2025-11-27 (234 days), 318 files, real_collection_present=True. The check reporting a true problem on day one. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…e-store feat(liveness): S3 — Appwrite production_forecasts check + real collection IDs (#241)
…_bucket (#242) python -m tools.liveness.unfao_delivery answers "when did FAO last receive anything?" per delivery stream (forecast_dataset_* / historical_dataset_*), with verdicts DELIVERING/STALLED/NEVER_DELIVERED per stream and an overall DELIVERING/DELIVERY_STALLED/UNREACHABLE/SKIP. TDD (12 tests first, fixtures = the real bucket listing); credentials reused from appwrite_store (same project/.env — S7 homes it in a shared module, noted); secrets never rendered; truthful SKIP without creds; exit 0/1/2. First live run: DELIVERY_STALLED — forecast 131 days (2026-03-10), historical 111 days (2026-03-30). The 4-month FAO stall (C-99-class silent lapse) is now machine-visible. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…elivery feat(liveness): S4 — FAO unfao_bucket delivery check (per-stream freshness) (#242)
…d a false idle verdict (#241, #242) Appwrite lists 25 files/page by default; sorting one page of the 318-file production_forecasts bucket yielded a FALSE 'newest = 2025-11-27' (truth: 2026-06-29 — the June monthly runs DID upload to the shelf). The shipped S3 check inherited the flaw; S4 was correct only because its bucket fits one page. Both checks now request server-side orderDesc($createdAt)+limit (S4: per-stream startsWith+orderDesc+limit, with per-stream totals); regression test pins the ordering requirement. Ground-truth docstrings corrected. Corrected live verdicts: production_forecasts STORE_ACTIVE (newest 2026-06-29, 19 days — July cycle NOT uploaded, a real finding); unfao DELIVERY_STALLED unchanged (131/111 days). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…ination fix(liveness): Appwrite server-side ordering — pagination bug gave a false idle verdict
…cycle? (#243) python -m tools.liveness.wandb_execution answers execution recency per monthly ensemble (roster mirrors monthly_run.sh, hand-encoded with a keep-in-sync note): latest finished forecasting run, created_at, state, days_since, and the train-window end month_id — the data-cutoff receipt that resolved the run-naming ambiguity. Verdicts COMPUTED/NOT_COMPUTED/ NEVER_RUN per ensemble; overall EXECUTION_CURRENT/STALE/UNREACHABLE/SKIP; exit 0/1/2. TDD (13 tests first, injected client/netrc/clock); wandb imported lazily (already installed, zero new deps); truthful SKIP without ~/.netrc api.wandb.ai. First live run: EXECUTION_CURRENT — pink_ponyclub/skinny_love finished 2026-06-29 (19d, cutoff 557), rude_boy/first_love 2026-07-15 (3d, cutoff 558). The 2026-07-19 trust-crisis question is now a one-command fact. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…xecution feat(liveness): S5 — wandb execution check (did the team compute this cycle?) (#243)
…244) python -m tools.liveness.vpn_store answers: does the legacy store hold a fresh fatalities run (computed uploads awaiting promotion)? Off-VPN it reports VPN_REQUIRED (exit 0) — the 2026-07-19 observation-boundary trap is now a named verdict, never a false RED. On-VPN: lists runs via views_forecasts.db_ops.ViewsMetadata, judges the newest fatalities run with the S1 parser + freshness budget (one parser, one convention). TDD (12 tests first); host-resolution failures classified VPN_REQUIRED, missing packages SKIP_NO_PACKAGE, else UNREACHABLE. Receipt encoded: the store's Postgres schema is literally 'forecasts_metadata' — the origin of the phantom Appwrite collection ID that killed the June 2026 run (legacy schema name copied into new-store config). The C-100 incident now has its full causal chain documented in code. On-VPN behavior pending live confirmation at the maintainer's next VPN session (noted in #244). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
feat(liveness): S6 — VPN store (gjoll) check with truthful VPN_REQUIRED (#244)
…#245) python -m tools.liveness runs all six surface checks, prints one raw-facts block per surface, and exits with the worst verdict (0 healthy/skips, 1 attention, 2 unreachable); a crashing check is contained and reported, never hiding the others. Extraction strictly limited to demonstrated duplication (WET-before-DRY, maintainer doctrine): report.py (fact renderer + the merged verdict->exit map, unknown verdicts fail loud) and appwrite_api.py (credentials dataclass/ .env parsing/resolution order, the pagination-cure query builders, the default stdlib fetch — previously duplicated across the two Appwrite checks). All six modules now delegate; appwrite_store re-exports for backward compatibility. Behavior preservation proven: all 92 pre-existing liveness tests pass UNMODIFIED; +10 runner/report tests (102 total). First full dashboard run: API fresh, input fresh (margin 6), shelf active (19d), execution current (4 ensembles), FAO stalled (131/111d — the one true red), vpn store VPN_REQUIRED. worst_exit: 1. The maintainer's stated bare-minimum MVP — "you must be able to check whether your forecasts are live" — exists. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
feat(liveness): S7 — DRY consolidation + python -m tools.liveness (the dashboard) (#245)
tools/liveness/README.md: one command + per-surface usage, the 0/1/2 exit contract (truthful skips are 0), every verdict per surface, and the three encoded conventions with receipts (data-cutoff run naming, the real Appwrite metadata IDs vs the phantom forecasts_metadata, server-side orderDesc listing). Validated against live output and the suite (102 passed); every command copy-paste correct. Register: C-100 -> Mitigated (the config-vs-reality exit exists as tools/liveness; residual = scheduling, which is C-99's exit); C-96 note (tripwire automated by datafactory_input); C-98 note (both stores now observable; authority still undecided); C-99 note (the instrument is the dead-man's-switch primitive; the heartbeat remains). Header 44/16. Closes the liveness epic #238. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
docs(liveness): S8 — README + register closure, C-100 → Mitigated (#246)
… charter coverage gap (T3) — falsify audit 2026-07-19 Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The 2026-07-19 falsification audit (claim: "tools.liveness is air and water tight") found 3 hard + 3 soft defects. Fixes: - report.one_line(): fact values with embedded newlines (live sqlalchemy errors) no longer break the one-fact-per-line contract (P1) - datafactory_input: missing datafactory_query is now a truthful SKIP_NO_PACKAGE (exit 0), mirroring vpn_store, not a false UNREACHABLE alarm (P2) - wandb_execution: _judge moved inside the per-ensemble try — a malformed created_at is a failure fact, never an uncaught crash (P4) - all six main()s classify the verdict BEFORE printing, so an unregistered verdict fails loud without a half-block the runner would contradict (P7) - roster tripwire: MONTHLY_ENSEMBLES is now asserted against monthly_run.sh itself, not a literal copy (P5) - README documents the C-102 non-goals (viewser, website, content sanity) so all-green cannot be over-read (P8, xfail-pinned) All findings enforced by tests/test_liveness_falsifications.py (107 liveness tests green, 1 xfail = the open C-102 gap). Register: C-101 -> Resolved, C-102 Open. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…-fixes fix(liveness): falsification-audit fixes — C-101 Resolved, C-102 registered
…p (T3) — falsify audit Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…coverage — C-103 Resolved Falsify audit (claim: "100% covered, green/beige/red all around") found the suite at 95% coverage, zero beige tests, and red misused for live network probes (ADR-005 red = adversarial/error-path — the maintainer's reading was correct; the suites' was not). - ADR-005 amended: fourth category `live` (real-external-service probes, truthful skip offline) + pyproject marker - 6 live probes relabeled red -> live; 22 genuine error-path tests red-marked - 8 beige structural tests added (module conventions, verdict-registration in the exit map, SURFACES registry, roster tripwire) - Coverage 95% -> 100% branch: real offline tests for the default clients (fake wandb/views_forecasts modules, monkeypatched urllib, credential resolution paths, netrc-probe failure paths); pragma only on __main__ guards, with reasons - tests/test_liveness_taxonomy.py enforces the taxonomy from now on 130 liveness tests green (was 107), 1 xfail (C-102 scope marker), ruff clean. Register: C-103 -> Resolved. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
test(liveness): ADR-005 'live' marker + taxonomy fix + 100% branch coverage (C-103)
…ing check
The serving probe sampled only /{run}/cm — a run serving country-month but
empty at grid level read LIVE_FRESH (the "one endpoint" gap, 2026-07-19).
Both public levels are now sampled at the first forecast month; a run empty
at EITHER level is LIVE_NOT_SERVING. Facts: serving_rows_cm +
serving_rows_pgm (replaces serving_rows_sampled); per-level errors prefixed.
Live receipt encoded: fatalities003_2026_05_t01 serves 2 rows at both
levels (cm: country_id keys, pgm: pg_id keys).
TDD (pgm-empty and both-levels-rendered tests written first); 132 liveness
tests green; branch coverage stays 100%.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
feat(liveness): old_api serving probe covers both cm and pgm levels
…d-model corrections + live blocker status for #230 Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
docs(reports): ADR-013 wire-contract review (views-models seat)
…gnment, liveness instrument (still Proposed; decision unchanged) Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
docs(adr-017): 2026-07-19 amendment — production evidence + ADR-013 alignment (stays Proposed)
…ion) + strengthen C-85 to Tier 2 falsify audit of the 2026-07-20 rusty_bucket delivery: four failed ensemble runs traced to two compounding silent defects. - C-85 raised Tier 3 -> 2: ensemble --saved reloads a stale y_pred.npy keyed on the sample-agnostic model-artifact timestamp; editing n_samples never changes the artifact, so the config change is silently discarded (no config fingerprint on the cache). Live evidence: mixed S=128/S=32 npy on disk, zero at the configured S=16; parent OOM ~18 GB every run regardless of config. - C-104 (new, Tier 2): posterior sample count is 4 different config keys (baseline n_samples, hydranet n_posterior_samples, r2darts num_samples, stepshifter pred_samples); the CI/parity contract reads only n_posterior_samples, a decoy for baseline — CI green while the runtime produces a different count. C-52's readiness fix seeded the decoy. The silence = three layers missing on one axis: no guardrail (cache/parity blind to config), fragmented knowledge (4 names + decoy), no test of the runtime pf.sample_count. Header Open 45 -> 46. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…nfig-load guard runs in CI, six tests off hand-patched paths (#470) * docs(#468): C-50 and C-42 Resolved — the packages they said were unpublished have been published C-50 said views-baseline was not on PyPI. It has been since 1.0.0 (2026-07-28); 1.0.2 (2026-09-09) is the first version that runs, all 37 models pin it (#460), and runtime_smoke.yml installs it from the index. C-42 said the synthetic models depended on an unreleased pipeline-core branch. That branch merged and released: PredictionFrameEnsembleManager imports from published 3.2.0 (verified), CI pins 3.0.1, the ensembles declare >=3.0.0 (#372). Both had been stale since July. Header: Open 70 -> 68, Resolved 43 -> 45. Tiers unchanged. reports/conda_to_uv_migration_investigation.md gets one dated line under its metadata: the three packages it proposes git+ dependencies for are on PyPI now. Body kept as the 2026-05-23 record it is. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01GG2cY2HaqmUMpwUyR5V82K * ci(#469): run the roster config-load guard in CI, in its own job tests/test_roster_configs_load.py constructs each of the 8 HydraNet roster configs through HydraNetConfig. It is the guard for C-259 — two production ensemble members sat unloadable from August to September while every test passed, because test_roster_conformance.py compares values and never builds the object. The load test importorskips views_hydranet, and run_tests.yml installs only pipeline-core, so it has skipped in CI since the day it was written. views-hydranet 0.1.0 reached PyPI on 2026-09-15 (views-hydranet#340); this job is where the guard now runs. Separate job, modelled on runtime_smoke.yml: run_tests.yml's install is pinned on purpose, and a red here means exactly one thing. #469 assumed torch was avoidable. It is not — importing config_initializer is torch-free, but get_config -> validate_loss_reg -> utils/utils.py:7 imports it; all 8 configs failed with ModuleNotFoundError when the job was first run locally. Installed from the CPU-only index, which is all a validator needs. pipeline-core and views-frames are not on this path and not installed. Released pin, not a git ref (C-106); ==0.1.0 moves with the models' ~=0.1.0 (#462). Verified locally in a venv built with exactly this job's install lines: 9 passed. Mutation — delete ss_feedback from heavy_freighter, where scheduled sampling is on — 2 failed, 7 passed: the load test on heavy_freighter and test_scheduled_sampling_declares_its_feedback. Reverted. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01GG2cY2HaqmUMpwUyR5V82K * test(#467): the six directory-creation tests use the _root redirect too Four ModelScaffoldBuilder tests and two EnsembleScaffoldBuilder tests set builder._model.model_dir and a hand-written builder._subdirs list after construction — the same incomplete-redirect shape #467 fixed for the script-building tests, and the module guard's own docstring named them as the second instance. They did not leak, because build_model_directory happens to read only the two attributes they patched. That is luck, not a property. Now: monkeypatch ModelPathManager._root to tmp_path before constructing the builder, and delete the hand-patching. The builder derives model_dir and _subdirs from the redirected root itself. EnsemblePathManager inherits _root, so the ensemble tests use the identical line. The assertions got stronger, not weaker. creates_subdirs and creates_gitkeep used to assert on a list the test invented (two paths, or one); they now assert on every subdirectory the path manager actually declares, and that each is under tmp_path. test_build_model_scripts_without_ directory_raises asserts the redirected directory does not exist before expecting the raise. test_build_model_directory_creates_dir took a monkeypatch fixture it never used; now it does. 25 passed; nothing written into the repository. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01GG2cY2HaqmUMpwUyR5V82K --------- Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
G4b status, 2026-09-15 — verified in substance; one key-scope re-run outstanding, and it is Simon'sThe gate says "unverified — needs FAO's caller key". That overstates what is missing. It was verified on 24 August. views-faoapi's own record ( Confirmed today by the faoapi session, read-only: Why it is not ticked yet: those calls and the 08-24 smoke ran on the write-scoped datastore key. faoapi's per-key cache (ADR-027) means a fresh read-scoped caller — which is what FAO is — hits a cold partition and can see a different state; that is the 2026-08-13 post-mortem exactly. So G4b should be ticked on a read-scoped run. No such key exists in the faoapi session's environment, and it correctly declined to look for one. What closes it: a read-scoped Appwrite key (the one FAO was issued, or one minted in the console), then from views-faoapi: Exit 0 → paste here → G4b met. Not an external-party dependency; a console action. Side finding: v1.7.2 and v1.7.3 are tagged and were never deployed. The served-code delta is one dict entry in the |
…-core 3.2.0, one file per source guard (#476) * test(#455): the guard #444 lacked — CI on 3.2.0, the sniffer test selects the file like the loader, no source may carry both PR #444 added config_maturity.py to 14 models and broke all 14 while 3,684 tests passed. Two reasons, both fixed here before a single config moves. CI pinned pipeline-core 3.0.1. Every 3.x before 3.2.0 crashes a config_maturity.py source at the sniffer — 3.2.0 (2026-09-08) is the first where the loader prefers the new file, deployment_status is no longer a mandatory key, and the five read sites accept either name (#495, #497). CI now runs 3.2.0, which the ensembles' >=3.0.0,<4.0.0 already permits. The pin's comment records the next bump's named trigger: from 3.3.0, get_queryset raises instead of returning None, and the catalogs job calls it for every model with no isolation. tests/test_core_config_sniffer_contract.py hardcoded config_deployment.py. It fed the sniffer the dict the real loader would have IGNORED whenever config_maturity.py existed, and stayed green. File selection is now delegated to pipeline-core's own load_maturity_config (3.2.0) — the same rule the managers use, so this test cannot drift from them again. _is_deprecated becomes _is_retired: 3.2.0 refuses to run a retired source by design, in either vocabulary. The #455 guard: tests/test_config_completeness.py::TestMaturityConfig asserts every source carries EXACTLY ONE of config_maturity.py / config_deployment.py, with the value valid in that file's vocabulary. Same for ensembles in test_ensemble_configs.py. ADR-017 Phase 2 is a rename; two files is the #444 state, where pipeline-core reads one and ignores the other with nothing noticing. config_deployment.py leaves the fixed required-file lists for the same reason: the requirement is "one of the two", not a filename. Mutation-tested on black_ranger, each reverted: both files present -> test_exactly_one_maturity_file FAILS, naming the source renamed, value "deployed" -> sniffer contract FAILS: "maturity='deployed' is not valid" renamed, value "retired" -> green; excluded from sniffer subjects, as 3.2.0 refuses it renamed, value "candidate" -> green — the dry run of the rename itself Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01GG2cY2HaqmUMpwUyR5V82K * feat(#449): every reader accepts both maturity vocabularies — the transition window, made real ADR-017 Phase 2 renames config_deployment.py -> config_maturity.py one source at a time, gated on the source's engine running pipeline-core >= 3.2.0. 69 of 119 sources (38 stepshifter, 31 r2darts2) run on 2.3.0, which requires the legacy file and knows nothing of maturity, so they keep it until their engine moves (views-stepshifter#103, views-r2darts2#24). During that window a source carries exactly one file, and it is the READERS that understand both. No config file moves in this commit; every reader is ready for the ones that will. deliveries/coherence.py — maturity_of() reads config_maturity.py first and returns the declared value after checking it is one of candidate/graduate/retired; falls through to the existing §3 translation for a legacy file; refuses a source carrying both. Its module docstring said "maturity does not exist yet". A latent bug in the error path is fixed on the way: it called _source_dir(source).relative_to(...), which is None for anything outside the repo. tools/catalogs/create_catalogs.py and update_readme.py — load either file; the catalog column is now "Maturity", in one vocabulary, translated for legacy sources. tests/test_tooling_scripts.py's "exact copy" of the table generator moves with it. run_integration_tests.sh — pre-flight classifies by maturity from whichever file exists; skips retired (or legacy deprecated) models as RETIRED. pipeline-core >= 3.2.0 refuses them anyway. tools/scaffold/build_{model,ensemble}_scaffold.py — new sources are born with config_maturity.py (candidate) via template_config_maturity. Verified: the builder writes the new file, not the old. tests/test_roster_configs_load.py _assemble, tests/test_delivery_coherence.py — prefer the new file; three new tests cover the declared path, the closed set, and the two-file refusal. Four CICs move with their scripts (cic-sync): CatalogExtractor, IntegrationTestRunner, ModelScaffoldBuilder, EnsembleScaffoldBuilder. WET, on purpose: the §3 table is now in coherence.py, create_catalogs.py and update_readme.py — three copies of four entries. deliveries/ runs standalone; the tools run under pipeline-core in CI; sharing an import couples them (coherence.py:71-83 says why). Revisit at a fourth reader. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01GG2cY2HaqmUBUwUyR5V82K * fix(#449): one reader of maturity, not three — catalog tools call coherence.maturity_of; ADR-017 §11 dated; docs follow the script Review findings on #476 (five agents + review-diff), all addressed: - tools/catalogs/create_catalogs.py, update_readme.py: the copied _LEGACY_TO_MATURITY dict mapped deployed -> graduate unconditionally, so the catalog showed ensembles/white_mustang as graduate while deliveries/coherence.py (R2: composite is graduate only if every member is) said candidate. Two readers of the same fact disagreed. Both tools now import deliveries.coherence.maturity_of; coherence.py is stdlib-only, so this couples nothing. The "WET, deliberately" justification in the plan was wrong and is withdrawn. - docs/ADRs/017_source_composition_delivery.md §11 Phase 2: the "cannot ride 3.0" blocker expired with pipeline-core 3.0.1/3.2.0; a dated status paragraph records that the rename is per source, gated on the engine's floor >=3.2.0, and that the window closes on views-models' say-so. Status line extended. - docs/run_integration_tests.md, README.md: the summary class the script now prints is RETIRED, not DEPRECATED. - run_tests.yml: the floor to bump with is the ensembles', not "the launchers'" (launchers do not pin pipeline-core). - two test docstrings: pipeline-core does not "silently" ignore the second file — it logs a warning and runs anyway. Full suite on 3.2.0: 7111 passed. ruff clean. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01GG2cY2HaqmUMpwUyR5V82K * docs(register): C-147 two maturity readers disagreed (mitigated), C-148 inert guard blind to dict readers (open); C-130 noted Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01GG2cY2HaqmUMpwUyR5V82K --------- Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
…rity.py — nothing graduates (#479) * feat(#449): rename the 50 sources on pipeline-core 3.x to config_maturity.py — nothing graduates ADR-017 Phase 2, the rename. 29 baseline, 8 hydranet and 13 ensembles — every source whose engine runs on pipeline-core 3.x — now carry configs/config_maturity.py -> get_maturity_config() -> {'maturity': ...}, generated from pipeline-core 3.2.0's own template so the fleet has one shape. The 69 stepshifter and r2darts2 sources keep config_deployment.py: their engines pin pipeline-core <3.0.0, and 2.3.0 requires the old file (views-stepshifter#103, views-r2darts2#24). Values, by ADR-017 §3: 40 shadow + 6 baseline -> candidate; 3 deprecated -> retired (bashful/dopey/happy_dwarf); the one deployed, ensembles/white_mustang, -> candidate under R2 (its members lavender_haze and blank_space are shadow). Nothing becomes graduate: a script is not an author, and maturity is the author's sign-off. With it: - tests/test_ensemble_maturity_rules.py replaces test_falsify_deployment_status_convention.py. That file's two xfail stubs said they would "flip to a hard gate when the rule is decided + the violation resolved"; ADR-017 §5 decided it and this rename resolved it. R1 and R2 are now fleet-wide hard gates in both vocabularies via deliveries.coherence.maturity_of (coherence.py's own R2 check only sees ensembles a delivery names, C-144). Mutation-tested: white_mustang -> graduate trips R2 naming both members; lavender_haze -> deprecated trips R1 for two ensembles. The second stub's subject ('production' literal in pipeline-core's ensemble check) was removed in pipeline-core 3.2.0 and no longer exists in any release CI installs. - tests/test_delivery_map_truth.py counts both files and both keys; the map states the figures per file (47 candidate / 3 retired across 50; 68 shadow / 1 deprecated across 69). Mutation-tested: a wrong figure in the map turns it red. - ADR-001, -003, -004, -009: the lines that stated the config_deployment.py contract, each with a dated status note (#453). README.md's two config descriptions rewritten. - KNOWN_REJECTED in test_core_config_sniffer_contract.py is unchanged: _is_retired already excludes the three retired sources from the sniffer's subjects, as the docstring says. - The root catalog and per-model READMEs regenerate from update_catalogs.yml on merge (it triggers on models/*/configs/config_*.py and commits); not hand-edited. Full suite on pipeline-core 3.2.0: 7112 passed, 0 failed, 0 xpass. ruff clean. validate_docs passes. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01GG2cY2HaqmUMpwUyR5V82K * docs(ADR-017): §11 status names #479 as the rename PR Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01GG2cY2HaqmUMpwUyR5V82K * docs(ADR-004): the vocabulary row is the maturity set; "production gating depends on it" was the false claim #453 named Review of #479 (agents 1, 3, 4 independently): the PR amended the required-keys row and left the next row asserting shadow/deployed/baseline/deprecated as the Tier-1 vocabulary with "production gating depends on it" — the exact line #453 called false. Now the closed maturity set, with the legacy four as the translated remainder on pipeline-core 2.x sources. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01GG2cY2HaqmUMpwUyR5V82K * docs(register): C-130 resolved by #479; C-144 narrowed — R1/R2 have a fleet-wide edit-time gate Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01GG2cY2HaqmUMpwUyR5V82K --------- Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
…d by the sniffer since 2026-06-27 (#482) #220 set evaluation_mode='point' on zero_/locf_/average_{cm,pgm}baseline and the three _dream fixtures without aggregate_method, which pipeline-core's CoreConfigSniffer has required for point mode since 2026-03-13. Every run of these nine died at config load, before any views-baseline code ran. One line each: "aggregate_method": "arithmetic_mean", the convention purple_alien and the hydranets already follow. Scope checked, not assumed: pipeline-core 3.2.0's sniffer over every non-retired model and ensemble refuses exactly these nine (tests/test_core_config_sniffer_contract.py); pipeline-core 2.3.0's sniffer — what the 68 non-retired stepshifter/r2darts2 sources actually run on — refuses none of them. KNOWN_REJECTED is therefore empty, as its comment always said it should become; removing one of the nine lines turns the contract test red naming the model. Proof of life on the versions requirements.txt installs (views-baseline 1.0.2, pipeline-core 3.2.0): diagonal_dream calibration --train --evaluate runs to "Done" — sniffer audited, model saved, metrics computed. Independent of #445/#459 (the targets outage), which is also closed. Claude-Session: https://claude.ai/code/session_01GG2cY2HaqmUMpwUyR5V82K Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
…odel whose data client is not installed (#483) * chore(#478): CI on pipeline-core 3.3.0; update_readme.py survives a model whose data client is not installed pipeline-core 3.3.0 (PyPI, 2026-09-19). The named trigger from #476's pin comment fired as written: the queryset loader now raises ImportError with an install hint instead of returning None (their #514), and tools/catalogs/update_readme.py called it for every model with no isolation — under 3.3.0 the catalogs job died on purple_alien, the first datafactory model, because the job installs no data client. Measured, not assumed: create_catalogs.py exit 0, update_readme.py exit 1 before; both exit 0 after, 23 models take the isolated branch, and the regenerated READMEs are byte-identical to what 3.2.0 produced (the None branch: "No description provided"). Rendering those querysets for real means installing views-datafactory in the job — that belongs with #474, not a pin bump. Guard: tests/test_update_readme_survives_missing_data_client.py runs the real script (it is monolithic, C-81/C-93) against a temporary repo holding one model whose config_queryset.py imports a client that cannot exist. Green on 3.2.0 and 3.3.0; red on 3.3.0 with the isolation removed. Both workflow pins 3.2.0 -> 3.3.0; the pin comment rewritten with what 3.3.0 changes for this repo (ADR-064 entity refusal, wandb <1.0, views-frames <3 — the CI venv now resolves wandb 0.30.0 and views-frames 2.0.0, the same resolution the ensembles get) and the next named trigger (4.0: viewser becomes an optional extra, their ADR-063). Full suite on 3.3.0 / wandb 0.30.0 / views-frames 2.0.0: 7257 passed, 0 failed. ruff clean. A real ensemble run on that resolution — synthetic_chant, calibration, --saved, after its three _dream members trained on views-baseline 1.0.2 — finishes with metrics. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01GG2cY2HaqmUMpwUyR5V82K * test(docstring): say what the missing-client guard covers and what C-81 still does not Review: the fixture's unknown client name takes the loader's bare re-raise path, not the install-hint path — deliberate, and now said; and the guard is narrower than C-81's validate=True crash at ModelPathManager construction, which stays open. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01GG2cY2HaqmUMpwUyR5V82K --------- Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
…<0.3.0 (#485) * build(deps): move all 31 r2darts2 models to views-r2darts2[manager]>=0.2.3,<0.3.0 r2darts2 0.2.3 (released 2026-09-19) is the first 0.2.x that installs beside the published platform: darts 0.40 / pandas <2 pins, wandb >=0.28.2 (pipeline-core 3.3.0 allows <1.0), plus fixes for three runtime bugs present on every 0.2.x before it. The `[manager]` extra is now required: on 0.2.x pipeline-core is optional in r2darts2, and every model's main.py imports it. Resolver-verified from PyPI alone (uv, py3.11, one model's requirements.txt): views-r2darts2 0.2.3, views-pipeline-core 3.3.0, viewser 6.6.4, darts 0.40.0, pandas 1.5.3, wandb 0.30.0. tests/test_requirements_hygiene.py and test_environment_sharing.py: 10 passed. 11 of these 31 models were run end to end on 0.2.3's code (editable) in the views_pipeline env before release; a fresh-env run on the published wheel is the check this PR still needs. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016d7duESygAoK9Dm9Vzwzs7 * test: the requirements-name parser stops at an extra — views-r2darts2[manager] is views-r2darts2 test_algorithm_coherence split the requirement on version operators only, so the [manager] extra every r2darts2 model now declares (#485) read as a different package from the one main.py imports. 318 passed. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01GG2cY2HaqmUMpwUyR5V82K * test(comments): the two blocks that said 0.2.x is not adoptable now say when and why it became adoptable test_requirements_hygiene.py and test_environment_sharing.py explained the <0.2.0 cap and the four conflicts behind it (#317). Both walls fell 2026-09-19 — pipeline-core 3.3.0 widened wandb, r2darts2 0.2.3 dropped darts 0.46 / pandas 2 — and #485 moves the spec. Dated, both states kept. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01GG2cY2HaqmUMpwUyR5V82K --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
… one point forecast of 31 that did not (#487) * fix(#486): bad_romance declares regression_point_metrics — the one point forecast of 31 that did not views-evaluation 2.0 refuses a point forecast with no regression_point_metrics ("No metrics configured for (regression, point)"), after training: bad_romance (num_samples 1, mc_dropout False) had the line commented out and cost an 816 s run on fimbulthul. Census of the 31 r2darts2 sources: it was the only one. Uncommented, in dancing_queen's shape. Guard: tests/test_point_forecast_declares_point_metrics.py — every source declaring num_samples <= 1 must declare regression_point_metrics. 29 point sources checked; red on bad_romance with the line re-commented. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01GG2cY2HaqmUMpwUyR5V82K * fix(#486): bad_romance's point baselines come back too — 8114034 commented out both lines, not one Review (git history): the commit that disabled regression_point_metrics disabled regression_point_baselines in the same hunk. dancing_queen — the sibling this fix says it matches — has both. Now it does. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01GG2cY2HaqmUMpwUyR5V82K --------- Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
#485 is not the pandas lock lifting Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01GG2cY2HaqmUMpwUyR5V82K
…202608 onto development (#491) * feat(#489): the eleven pgm datafactory r2darts2 models, from staging_202608 onto development Story S2 of epic #488. blue_ocean brave_heart dancing_monkey dark_necessities dark_river golden_eagle little_talks mister_bluesky old_rules red_hawk silent_fox — NBEATS / TSMixer / TiDE, pgm, all three targets, datafactory (loa priogrid_month). Copied directory by directory from origin/staging_202608 (no branch merge; the four staging ensembles over them stay on #402), then four corrections each, applied by a one-shot script that is not committed: - requirements.txt: `views-r2darts2>=0.1.0` (unbounded, no client) -> the fleet spec `views-r2darts2[manager]>=0.2.3,<0.3.0` plus `views-datafactory>=1.9.0,<2.0.0`, which the queryset imports. - run.sh: env_path pointed at envs/views-hydranet — a copy-paste from the hydranet template — now envs/views_r2darts2 like the 31 siblings. - config_meta.py: prediction_format "prediction_frame" -> "dataframe". Every other r2darts2 model runs on "dataframe" (13 of them passed on fimbulthul today); staging hid these 11 from test_pfe_production_readiness.py with a conftest skip-list that development does not have, and they declare neither n_posterior_samples nor hp regression_targets. The prediction_frame migration is a separate matter. - config_deployment.py (shadow) -> config_maturity.py (candidate), from pipeline-core's template: on 3.x from day one, one file per source. New guard, tests/test_datafactory_client_is_declared.py: a config_queryset.py that imports datafactory_query must have views-datafactory in requirements.txt — the state the eleven arrived in, which no existing hygiene test judged. 34 importers, 34 declare; red when one is dropped. Counts that moved: test_environment_sharing views_r2darts2 22 -> 33; forecast_delivery_map 61 config_maturity.py (58 candidate / 3 retired), 130 files, 131 source directories. Full suite on pipeline-core 3.3.0: 7682 passed. The sniffer contract accepts all eleven (126 subjects, none refused). ruff and validate_docs clean. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01GG2cY2HaqmUMpwUyR5V82K * docs(#489): the eleven READMEs regenerated on 3.3.0 — maturity, tree and targets as the configs say; comment count 31 -> 42 Review: the copied READMEs still said 'shadow' and listed config_deployment.py; dancing_monkey's listed one target of three. Regenerated with tools/catalogs/update_readme.py under the CI mirror — the bytes the merge job would have written — so nothing on development is stale for a minute. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01GG2cY2HaqmUMpwUyR5V82K * docs(register): C-149 — run.sh's env prefix is a prompt answer nothing checks against requirements.txt Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01GG2cY2HaqmUMpwUyR5V82K --------- Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
…: True on views-hydranet 0.1.1 (#493) * feat(#484): the eight HydraNets refuse to train on CPU — require_cuda: True on views-hydranet 0.1.1 Story S3 of epic #488. views-hydranet 0.1.1 (2026-09-19) adds `require_cuda`: True turns a CPU device into a RuntimeError at the top of training, evaluation and forecasting instead of a warning banner. The reason is views-hydranet#377: a fresh install resolved a torch whose CUDA build the machine's driver could not run, torch fell back to CPU, and violet_visitor trained for 6 h 46 m at 1/50th speed under a banner. On a server, that is how a week is lost. - 8 x configs/config_hyperparameters.py: 'require_cuda': True, next to freeze_recurrent. - 8 x requirements.txt: views-hydranet~=0.1.0 -> ~=0.1.1. Not optional: HydraNetConfig has extra="allow", so on 0.1.0 the key is accepted and does nothing — the floor and the key land together or the key is decoration. - .github/workflows/roster_configs_load.yml: 0.1.1. - tests/test_roster_conformance.py FOUNDATION: require_cuda True, pinned per model. Measured: - roster tests (test_roster_configs_load constructs all 8 HydraNetConfigs, conformance) on a venv built as the CI job builds it — pipeline-core 3.3.0, CPU torch, views-hydranet 0.1.1 --no-deps: 201 passed. Full suite on 3.3.0: 7726 passed. - The point of the story, end-to-end: a fresh venv from violet_visitor's own requirements.txt with torch swapped for the CPU build (the #377 shape), `main.py -r calibration -t`: config audited, pgm data fetched and audited, then "require_cuda=True but the device for training is 'cpu' ... Refusing to run on CPU (#377)" — exit 1 after 17 s, zero epochs. - Mutation: remove the key from bold_comet -> test_foundation_value red, naming it. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01GG2cY2HaqmUMpwUyR5V82K * ci(comment): the roster job's header names the pin it must match by date, not by a literal that just went stale Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01GG2cY2HaqmUMpwUyR5V82K * docs(register): C-150 — a fresh env's torch can outrun the driver; loud since #493, fixed by nobody yet (#494) Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01GG2cY2HaqmUMpwUyR5V82K --------- Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
… red on the header count (#495) The merge-readiness test counts '| Status | Open |' exactly; 'Open — made loud, not fixed' was not counted and development went red on the #493 squash. Qualifier lives in Notes, where it was already said. Claude-Session: https://claude.ai/code/session_01GG2cY2HaqmUMpwUyR5V82K Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
…he 38 stepshifter sources remain on the legacy vocabulary (#496) * feat(#490): config_maturity.py for the 31 cm r2darts2 models — 38 stepshifter sources are all that is left on the legacy vocabulary Story S4 of epic #488; ADR-017 Phase 2, the r2darts2 family. #479 could not rename these 31: their engine pinned pipeline-core <3.0.0, and 2.3.0 requires config_deployment.py. #485 moved the engine (views-r2darts2 0.2.3, pipeline-core >=3.0), so the #479 procedure applies: each file generated from pipeline-core 3.3.0's template_config_maturity, values by ADR-017 §3 — all 31 were `shadow`, all 31 are `candidate`. Nothing graduates; a script is not an author. One shape across the fleet (the 92 new-vocabulary files hash identically modulo the value). With it: docs/forecast_delivery_map.md restated — 92 config_maturity.py (89 candidate / 3 retired), 38 config_deployment.py (37 shadow / 1 deprecated), 130 files, 131 source directories; ADR-017 §11 status: the window now closes when views-stepshifter#103 ships. test_delivery_map_truth follows the figures; KNOWN_REJECTED stays empty; the catalogs job regenerates the 31 READMEs on merge (it did so for #479 and #491). Full suite on pipeline-core 3.3.0: 7725 passed — the sniffer contract accepts all 31 through load_maturity_config, and R1/R2 (test_ensemble_maturity_rules) hold. ruff and validate_docs clean. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01GG2cY2HaqmUMpwUyR5V82K * docs: README and ADR-017's status line no longer call r2darts2 a legacy-vocabulary engine Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01GG2cY2HaqmUMpwUyR5V82K * docs: the maturity gate is the resolved pipeline-core, not the declared floor; README names only the engine still on 2.x Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01GG2cY2HaqmUMpwUyR5V82K --------- Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
…iewser | datafactory | synthetic (#497) * feat(#474): the catalog says which data source each model reaches — viewser | datafactory | synthetic Story S5 of epic #488. #473 measured 77 models on viewser (pandas 1, Python <=3.11) and 29 on the datafactory or synthetic data, and nothing showed it: the queryset file knew, the catalog and READMEs did not. Now: - tools/catalogs/data_source.py — one reader, by AST, never a guess: `viewser` (imports viewser), `datafactory` (imports datafactory_query), `synthetic` (generate() returns {"source": "synthetic"}), `none` (no file), `unknown` (both clients, or neither). Reads the source; imports nothing from it, so it answers where the client is not installed (#483). The same shape as deliveries.coherence.maturity_of (C-147). - create_catalogs.py: a `Data Source` column in the root README's model table, after Input Features. update_readme.py: a `Data Source` row in every model README. Both import the one reader. - models/README_scaffold.md, ensembles/README_ensemble_scaffold.md: the row that has held the maturity value since #476 is now labelled Maturity, not Deployment Status; models get the Data Source row. - READMEs regenerated here (root + 111 models + 10 ensembles): the catalogs job triggers on models/*/configs/config_*.py, which this change does not touch, so the output moves with the script in the same PR. - tests/test_data_source_catalog.py: each classifier branch on a synthetic file (viewser, datafactory, synthetic, both -> unknown, neither -> unknown, relative import -> unknown, no file -> none); no model in the fleet is `unknown`; and the pin — 77 viewser / 34 datafactory / 6 synthetic of 117 — changed on purpose when a model migrates. test_tooling_scripts's characterisation copy follows the generator. CatalogExtractor CIC lists the key. Full suite on pipeline-core 3.3.0: 7735 passed. ruff and validate_docs clean. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01GG2cY2HaqmUMpwUyR5V82K * review(#474): the synthetic check reads generate()'s return only; the fleet test imports conftest's discovery; docstring and CIC name the key Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01GG2cY2HaqmUMpwUyR5V82K --------- Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
…ple_alien's exclusion was an env-isolation workaround (#500) * chore(#499): the integration runner excludes nothing by default — purple_alien's exclusion dates from when the HydraNets could not run The default was set while purple_alien (a HydraNet) was unrunnable; all eight run now (#488). The banner and summary print 'none' for an empty list; the CIC and the user doc say when and why the default went. --exclude still replaces rather than appends. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01GG2cY2HaqmUMpwUyR5V82K * docs(#499): the purple_alien default was an env-isolation workaround (5a2fd2e, 2026-03-15), not a broken model — say so; README's flag table follows Review: git history says the runner moved to one shared conda env and purple_alien alone needed views-hydranet, which that env lacked; it was the only HydraNet in the tree then. The 'could not run' wording was mine and wrong. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01GG2cY2HaqmUMpwUyR5V82K --------- Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
…r the first integration pass (#501) * feat(#499): the eight HydraNets read global land; total_lessons 40 for the first integration pass Runbook #499 Step 1. REGION "africa_me_legacy" (13,110 cells) -> "land" (64,818 — the set the FAO delivery cuts land_gaul from) in all eight config_queryset.py. total_lessons 300 -> 40 in all eight hyperparameter files, commented as what it is: a run-time budget for the pass that measures whether a HydraNet at global land fits fimbulthul, not a model choice; 300 (#463) is restored by a follow-up PR once the pass has run. total_lessons is deliberately unpinned by the roster tests for exactly this reason. The roster test that pinned the region asserted `"africa_me_legacy" in <file text>` — a substring, which the comment recording the history now satisfies, so it passed vacuously on the new region. Replaced by test_global_land_region, which reads the REGION assignment line and pins "land". Mutation: one model back to africa_me_legacy -> red, naming it. Not changed: the FAO postprocessor's own queryset (still africa_me_legacy in git; #499 B5, C-110), and every other config value on the eight. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01GG2cY2HaqmUMpwUyR5V82K * review(#499): the lessons comment names its trigger and owner; violet_visitor's docstring stops saying africa_me_legacy Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01GG2cY2HaqmUMpwUyR5V82K * docs(register): C-146 — a third occurrence: the region guard matched a substring, not the assignment (#501) Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01GG2cY2HaqmUMpwUyR5V82K --------- Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
…factory fetch on current Python 3.11 (#502) Claude-Session: https://claude.ai/code/session_01GG2cY2HaqmUMpwUyR5V82K Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
…, offsets 1/1 — global land does not fit the Africa+ME crop (#503) * feat(#499): the eight HydraNets' volume is the whole PRIO-GRID raster — global land does not fit the Africa+ME crop Runbook #499, found by the first pass on fimbulthul (2026-09-19 21:42): with REGION "land" (#501) views-hydranet's DataSniffer refused every model in 9 s — "Spatial Span Violation! Data spans 278.0x719.0, but volume resolution is 180x180". The four crop values row_offset 87, col_offset 310, height 180, width 180 are the Africa+ME window. The sniffer's own docstring names the global shape ("region=land in a 360x720 volume ... leading rows/cols zero, expected"): row_offset 0, col_offset 0, height 360, width 720, on all eight — the roster test requires the eight to share one grid topology, and they do. Not measured here: memory. The volume is 8x the crop's area; training samples 32x32 windows (window_dim), so the per-step cost need not scale with it, and views-hydranet's disk_guard refuses an oversize posterior cube before allocating. The pass measures it. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01GG2cY2HaqmUMpwUyR5V82K * fix(#499): row/col offsets are 1, not 0 — pipeline-core numbers the grid from 1 (the April fix, 9540a7b, applied to all eight) Review: pipeline-core 3.3.0 dataloaders.py:540 builds row = (pgid-1)//720 + 1, so rows run 1..360 and cols 1..720; views-hydranet's VolumeHandler indexes row - row_offset and refuses r_idx >= height. With 0/0 the last row/column of land overflows the 360x720 volume — or, if no land cell sits on the edge, every prediction lands one cell off. 9540a7b set 1/1 for heavy_freighter on this exact combination in April; f0b4436 later returned it to the crop. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01GG2cY2HaqmUMpwUyR5V82K --------- Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
… — their predictions were labelled country_id and refused (#504) * fix(#499): the eleven pgm darts models declare entity_id: priogrid_id — their predictions were labelled country_id and refused Runbook #499. dark_river's calibration run on fimbulthul (2026-09-22) trained on 64,818 priogrid entities, then had all 13 evaluation sequences refused by pipeline-core's CorePredictionSniffer: MultiIndex names ('month_id', 'country_id') do not match expected layout for level='pgm': ('priogrid_id', 'month_id') views-r2darts2 reads the index name from the combined config and defaults it to "country_id" whatever the level (darts_forecasting_model_manager.py:140, published 0.2.3) — right for the 31 cm models, wrong for these 11. The prediction VALUES were always priogrid cells; only the label was wrong. Filed upstream as views-r2darts2#55 (derive it from level, as dataset/base.py:934 already does); this declares it explicitly until that ships. Nobody saw the refusal because pipeline-core evaluates in threads and discards their exceptions: zero predictions written, run reported PASS (views-pipeline-core#529). That is how three days of GPU produced metric frames and no data. Guard: tests/test_darts_entity_id_matches_level.py — every r2darts2 model's effective entity_id (declared, else the engine's country_id default) must match its level. 43 checked; red when one of the eleven drops the key. Full suite on pipeline-core 3.3.0: 7773 passed. ruff clean. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01GG2cY2HaqmUMpwUyR5V82K * review(#499): the level lookup fails readably, and the third 'is it r2darts2' discovery names its trigger Review of #504: EXPECTED[level] would have raised a bare KeyError if the platform ever gained a third level (today impossible — test_config_completeness pins cm|pgm); now it says so. And this is the third way the suite asks whether a model is r2darts2 — left as three, with the fourth caller named as the trigger to extract one helper into conftest. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01GG2cY2HaqmUMpwUyR5V82K --------- Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
…hers consume (#506) * feat(#505): collapse posterior draws to the point predictions researchers consume (#506) The eight HydraNets emit D x K draws per cell, not a number. `ensemble-updater` takes one number per cell. Nothing currently bridges that. The platform already defines the bridge: `inference_orchestrator.py:175` inverts the log1p scaling, `:179` collapses the draw axis by the model's declared `aggregate_method` (views-hydranet's ADR 021 / ADR 039 stage 5). Stage 5 is gated on `evaluation_mode == "point"` and all eight run `stochastic`, so the draws reach disk uncollapsed and nothing downstream folds them. `tools/collapse` performs that stage, on the operator's machine, from the preserved numpy. It honours the declared method rather than hard-coding one — the first draft hard-coded the mean, which is the "two places to state one fact" pattern vmo_021 exists to stop. `test_collapse_declaration_matches_the_converter` pins the converter's default against all eight models' `aggregate_method`, and against their `evaluation_mode`, so the two cannot drift apart silently. Verification: - 28 tests: contract, input mutation, and real output already on this machine. - 21 mutations applied to the converter, 21 caught, 0 survived. The first pass caught 19; the two survivors (lexicographic origin ordering, and a dropped identifier-length check) were real coverage holes, which is why the fixture now carries 13 origins rather than 3. - Cross-checked against a pure-Python per-row recomputation over all 13 origins of all eight models' validation output: zero mismatches. - Testing found two defects in the converter: a guard that raised ValueError while building its own error message, and a float32 accumulator. Not validated: global land, the calibration partition, D x K = 32. Everything measured here is Africa+ME validation at 16 draws, because that is what exists locally. ADR-023 says so in Validation & Monitoring. ADR-023 records the decision; C-152 records the coupling it accepts (a third reader of an on-disk layout no repo publishes). Both 021s are cited, so the README's cross-repo prefix rule now covers vhy_021 / vmo_021. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01GG2cY2HaqmUMpwUyR5V82K * fix(#505): the roster test imports tools.collapse lazily — roster-configs-load has no pandas `roster-configs-load` installs views-hydranet and nothing else, then imports tests/test_roster_conformance.py through tests/test_roster_configs_load.py. A module-level `from tools.collapse.collapse_predictions import ...` pulled pandas into that job's collection and failed it at import, before any test ran. The import now sits inside the one test that uses it, with a comment saying why it must stay there. Verified by importing the module with pandas blocked from sys.meta_path. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01GG2cY2HaqmUMpwUyR5V82K * fix(#505): the scale guard could not fire on a single corrupted target Code review (5 agents) found the one defect that mattered, and it was a guard that could not fire. `_check_scale` took the maximum of the three targets FLATTENED TOGETHER. A scaler-registry mismatch upstream does not have to hit all three targets, and when it hits one, a healthy sibling carries the combined maximum over the threshold. The corrupted target then ships as log1p(count) — a plausible-looking parquet with a silently wrong column, which is the exact failure this guard was written to prevent. Its test could not detect the difference either: it set all three targets to log-space values at once, so it passed under the broken implementation and the correct one alike. A test that cannot distinguish the two designs is not evidence about either. - the check is now per target, and names the offending column - new test puts ONE target in log space with healthy siblings; a mutation that restores the flattened form is now caught (M22) - added a duplicate (month_id, priogrid_id) guard: ensemble-updater joins on that pair, and every target agreeing on a duplicated identifier is still a duplicate, so cross-target alignment is blind to it (M23) Also from review: - module docstring named `--max-plausible`; the flag is `--min-plausible-max` - bare "ADR-069" cross-repo citation qualified to views-hydranet `vhy_069`, per the README's own prefix rule - the docstring described only the mean; the module honours `median` too - docs/ADRs/README.md listed `[ADR-021]` while the paragraph directly below it documents that 021 collides across repos — now `[vmo_021]`, applying the convention the PR itself introduces - `test_collapse_declaration_matches_the_converter` moved out of TestGridAndTargets (documented as "Region grid and target channels") into its own class Verified: 8037 passed, ruff clean, 23/23 mutations caught with 0 survivors and 0 unrun, tree byte-identical after the campaign. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01GG2cY2HaqmUMpwUyR5V82K --------- Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
…ation run (#507) * feat(#499): the eight HydraNets train 300 lessons for the real calibration run 40 was the integration-pass value set by #501 to make the first global-land pass cheap. It is not a training length: measured on a rented RTX PRO 4500 SE, one lesson is 84 seconds, and the bias against observed fatalities is roughly twice as bad at 160 lessons as at 300 (0.05 vs 0.13 predicted/observed total, measured on the validation output already on the operator's laptop). Cost of this line, measured rather than guessed: 300 lessons is ~7 h of training plus ~1 h of evaluation per model, so ~8 h and ~$6 per model at $0.75/hr, and ~$48 for all eight. `n_head_samples` deliberately NOT raised. The plan had been D x K = 4 x 8, but `ensembles/rusty_bucket/configs/config_hyperparameters.py:16` pins `expected_samples_per_model: 16`, and its own comment records that the pool was thinned from 1024 to 128 because the full-S run peaks at ~28.6 GB. Raising K to 8 would break that pin, double each model's prediction volume from 13 GB to 26 GB, and lengthen evaluation — to buy finer quantile resolution that the estimator analysis did not need at 16 draws. If the ensemble's memory behaviour is fixed later, K and `expected_samples_per_model` move together, in one PR, not this one. Verified: full suite green (7997 passed), roster conformance, ensemble configs and the ADR-015 sample-count report all unaffected, since only the lesson count moves. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01GG2cY2HaqmUMpwUyR5V82K * fix(#499): the comment above total_lessons said the restoring PR was still owed Code review (5 agents) on #507. Two findings, both about text that would mislead the next reader rather than about the value itself. 1. The comment directly above the changed line, identical in all eight files, read: "40 for the eleven-model calibration pass ... Production is 300 (#463); the PR that restores it is owed by whoever ticks that pass on #499." After this change it lied twice over — it named a value the line no longer holds, and it described as owed the very work it sits on top of. Rewritten to say what is true now, and to record the measured cost (84 s per lesson at global land, so ~7 h per model) so the number is not a mystery to whoever reads it next. The stale "eleven-model" phrasing, which belonged to the darts group and not to these eight, goes with it. 2. `run_integration_tests.sh`'s 1800 s default was sized when these models trained 40 lessons. At 300 they need ~7 h and WILL report TIMEOUT. That is a budget, not a regression, but a person reading a wall of TIMEOUT rows would reasonably conclude otherwise. The header and --help now say so and give the flag. Comments only; no behaviour changed. Two corrections to my own claims in the PR description, both raised by review: - I called the quality evidence "measured, not guessed". It is weaker than that. The 0.05-vs-0.13 predicted/observed figures compare 160-lesson and 300-lesson models on Africa+ME VALIDATION output, not the 40-vs-300 before/after this diff actually makes, and not at global land. No 40-vs-300 ablation exists. It is a reasonable proxy — 40 was never claimed to be adequate — but the framing overstated how directly it was measured. The 84 s per lesson figure IS directly measured, on the pod, and is the only number here I would defend as such. - The gate the code comment named was the eleven-model fimbulthul pass, which never formally closed. The evidence offered instead is a fresh end-to-end pod run. Same intent — prove it fits, measure the cost — but different evidence than the gate specified, and that substitution should be visible rather than glossed. Also merges origin/development (#506, the converter) so the branch is current and the suite runs against the merged state. Verified: 8037 passed, ruff clean, bash -n on the runner. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01GG2cY2HaqmUMpwUyR5V82K * docs(#499): the runner's CIC records that a HydraNet TIMEOUT is now the budget `cic-sync` was right to fail. `run_integration_tests.sh` is CIC-governed by docs/CICs/IntegrationTestRunner.md (ADR-006), and my change there was comments only — so `[cic-skip]` in the PR title would have passed the gate. That would have been the wrong way through it. The contract document is where someone looks to find out what the runner guarantees, and "the default timeout no longer fits eight of the models in the roster" is exactly that kind of fact. Put in the title, it would have been invisible a week later. So the CIC now carries it: why 1800 s was right at 40 lessons, why 300 lessons means ~7 h of training plus ~1 h of evaluation per model, the flag to use, and why the default is deliberately NOT raised — it suits every other library, and raising it globally would turn a genuine hang in a cheap model into a half-day wait. The two tables that a reader hits first — the `--timeout` row and the "Model exceeds timeout" failure-mode row — now both point at that note, because a wall of TIMEOUT rows is the exact shape a real failure takes and nobody should have to guess which one they are looking at. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01GG2cY2HaqmUMpwUyR5V82K --------- Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
…nsit, and we put it on rented hardware (#510) The datafactory speaks plain HTTP. HTTP Basic sends the credential base64-encoded on every chunk request, and base64 is encoding, not encryption. views-datafactory accepted that as their C-318 when the audience was a trusted circle on trusted networks, which on fimbulthul was reasonable. On 2026-09-28 we changed the audience without changing the mechanism: the credential went onto five rented machines in datacentres we do not control. The operator chose knowingly to use his personal login rather than provision a throwaway, having been told what it meant. That is recorded, not second-guessed — the work was owed and the server was gone. What makes the residual risk outlive the run is two properties of the credential itself: no expiry and no per-host registration. It stays valid until a person rotates it by hand, and it authenticates from anywhere. So a pod image, a snapshot or a detached volume that survives a campaign carries a live permanent credential, and only housekeeping closes that. Tier 3: a real but unquantified interception risk rather than a demonstrated compromise, with cheap known mitigations. The trigger is deliberately the general case — any run on hardware we do not own — not RunPod specifically, because the next one may be a collaborator's machine or a CI runner. Mitigations in preference order: a throwaway login retired at campaign end (about three commands for whoever administers the data server); TLS on the data server, which removes the class entirely; or deleting rented volumes and images at the end and rotating afterwards. Raised by the views-datafactory session reviewing the RunPod post-mortem (#508). Cross-refs C-151, views-datafactory C-318, and #509 — the client floor still permits a version that carried this credential across redirects to other hosts. Header: 160 -> 161 entries, 151 -> 152 concerns, Accepted 4 -> 5, T3 64 -> 65. Verified: 8036 passed, register header tests green. Claude-Session: https://claude.ai/code/session_01GG2cY2HaqmUMpwUyR5V82K Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
…rator guide (#508) * docs(#499): the first RunPod deployment — post-mortem, cost note, operator guide Three documents from one effort, split by audience because they were fighting each other in a single file. Plus `tools/podrun` v0.1.0, and a correction to a number that #507 merged. - reports/postmortem_runpod_first_deployment_2026-09.md — what happened and what must change. Owns the narrative, the misdiagnoses, and the failures. - reports/runpod_cost_and_time_note_2026-09.md — what it cost. Written for a manager, cost and time only, one clearly-labelled extrapolation. - docs/runpod_run_guide.md — how to do it. Owns the procedure and the gates. - tools/podrun/ — the runner, marked v0.1.0 PROVISIONAL in four places, with the conditions for promoting it out of 0.1.0 stated. Each fact has one owner; the others cross-reference. The guide is the how, the post-mortem is the why, per fao_delivery_runbook.md's precedent. THE CORRECTION. #507 merged "one lesson is 84 s, so 300 lessons is ~7 h" into the eight config comments, run_integration_tests.sh and the runner's CIC, stated as measured fact. It is wrong. The 84 s came from the FIRST lesson of a cold two-lesson smoke run, which carries warm-up and is not representative of the other 299. Three completed models give 202, 253 and 272 minutes end to end — 40-54 s per lesson, ~4 h per model. Corrected in all three places, with the error named rather than quietly overwritten. No operational harm: the recommended --timeout 30000 was over-provisioned against the wrong number and remains safe against the right one. This repeats, within hours, the lesson the post-mortem's own §2.6 records — never characterise against a throwaway model's output. It is now §7 item 8. REVIEW. Three peer sessions reviewed independently and each found something real. views-hydranet: the guide stated two different runtimes (the above); the disk_guard problem is worse than described — even taught to read the cgroup it counts only the posterior cube, missing the ~2.3 GB input volume and the torch context, so it would still under-report peak by 2-3x; the 2 steps/s threshold is pod-relative, not a hardware expectation (a 4070 laptop does ~15); both q95 figures were measured at S=16, so the claim is stability across datasets, NOT across sample counts; and C-151 must be cited as views-models C-151 because views-hydranet's C-151 is an unrelated entry. views-faoapi: "the upload can run from a machine Simon controls" was a procedure claim I was not entitled to — the architecture supports it (UPLOAD_ENABLED defaults False, no store client constructed when disarmed) but there is one entrypoint, no --no-upload, and disarming edits a committed delivery declaration. Restated as architecture + open question. Also: the ensemble and delivery steps are per-run and memory-bound, so the per-model extrapolation is silent about them; and the two documents disagreed on how many models were complete. views-datafactory: the preflight check I wrote passes only because urllib follows a 308 and carries the Authorization header across it — the exact pattern their #388 removed. Verified on a live pod: bare URL 308, /.zmetadata 200, and Authorization sits in req.headers not unredirected_hdrs. Now points at .zmetadata. Also raised the client floor to >=1.13.0, where the credential-handling fixes landed. Verified: 8036 passed 0 failed, ruff clean, validate_docs.sh passes, test_tools_layout passes with the new group. Downloaded predictions deliberately not committed. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01GG2cY2HaqmUMpwUyR5V82K * fix(podrun): the draws archive could be empty and still report success Bug review of tools/podrun/pod_run_model.sh before merging #508. Eight successful runs had exercised the happy path; this is about the paths they did not. CRITICAL, and verified empirically rather than reasoned: find <no matches> -print0 | tar --null -T - -> exit 0, valid 22-byte archive So compress_draws could write an empty posterior and the script would proceed to STATUS=OK. The only other signal was a small number in MANIFEST that a human had to notice. On the one model family this has run the find pattern matches; on the next one it might not, and `__init__.py` already warns that no other family has been through it. Now counted before and verified after: - refuse if find matches zero lr_* draw files - after tar, list the archive and require the entry count to equal the count found HIGH: two config guards checked file TEXT, not the parsed value. A regex took the first `'total_lessons': <digits>` anywhere in the file and a grep matched `REGION = "land"` anywhere including comments. That is the defect this repo shipped once already — #501's "the guard that was not one", where a substring assertion was satisfied by a comment recording the region's history. Both now import the config module and read the value the interpreter sees. Verified: reads 300 and "land" from violet_visitor, and the file does contain "300 lessons" in a comment, which is the surface a text check exposes. HIGH: `SRC=$(ls -d ...predictions_calibration_*)` had no existence guard. An empty result meant `cd "$SRC/.."` -> `cd /..`, i.e. the filesystem root. It happened to fail loud because find errors on an empty path argument, which is an accident to depend on. Now refuses explicitly. MEDIUM: STATUS was never cleared at start, so a stale FAILED from an earlier attempt shadowed a fresh run for its whole multi-hour duration -- and the script's own header advertises `cat STATUS` as the way to watch from outside. Cleared now. MEDIUM: output directories were not cleared between attempts. Parquets are named from the SOURCE run's timestamp, so an old set and a new set could coexist; if they summed to 13 the count check would pass while the manifest covered two different training runs. Both output dirs are now cleared per attempt. MEDIUM: the model and converter existence checks had no stage of their own, so die() reported whatever stage ran last -- FAILED:verify_env for "no such model". MEDIUM: die()'s explanation went only through the tee subshell, which can lose its last buffered lines on a hard kill, at exactly the moment the reason is wanted. The reason is now also written directly to $OUT/FAILURE, as STATUS and STAGE already were. LOW-MEDIUM: no lock, so two invocations for the same model on one pod raced on the clone, the log and STATUS. I nearly caused this myself today when a chain and a manual assignment could both have taken bright_starship. Now takes $OUT/.lock and releases it on exit. Known and NOT fixed, recorded in __init__.py and the post-mortem's open items: the script has no automated tests, and MANIFEST's git sha and lesson count are unchecked, so they would ship blank rather than refuse. Both are accepted at v0.1.0. Verified: all three guards fire on constructed inputs (empty archive refused; counts match on real files; empty SRC refused), the config check reads real values through the interpreter, bash -n clean, ruff clean, test_tools_layout passes. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01GG2cY2HaqmUMpwUyR5V82K --------- Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
…ers (#522) * fix(#439): bump VIEWS_POSTPROCESSING_PIN 1.1.1 -> 1.4.0 in both launchers Closes #439. Two lines, one value, and it is the only thing standing between the FAO findability fix and a pod. WHY THIS IS NOT A NO-OP. The launcher does NOT install views-postprocessing from PyPI. It installs from a git ref: tools/launcher/postprocessor.sh:132 pip install "git+https://${GITHUB_TOKEN}@.../views-postprocessing.git@${VIEWS_POSTPROCESSING_PIN}" and refuses the run at :164 if the installed ref does not match the pin. So views-postprocessing#314 being merged does nothing for a pod, and releasing it to PyPI (1.4.0 is there) does nothing either. Verified empirically rather than reasoned: the 2026-09-29 pod ended up with views-postprocessing 1.1.1 while PyPI was at 1.3.0. This was invisible because the two halves of the same incident reach a pod by DIFFERENT mechanisms — views-pipeline-core transitively from PyPI through views-hydranet's dependency range, views-postprocessing from a git tag by exact pin. "Is the fix released?" therefore has two different correct answers depending on which half is meant, and nothing states that anywhere. views-postprocessing#310 is the same family. WHAT 1.4.0 CARRIES: #314, the C-94 findability guard that verifies every artefact by name rather than trusting a returned file id — the defect that let the 2026-09-29 delivery report "Postprocessor Run Completed" while being unservable; and #315, a guard deriving the port's documented datastore contract from source so the prose cannot drift from the code. VERIFIED BEFORE COMMITTING, not assumed: - tag 1.4.0 resolves on the remote the launcher actually pulls from (8db8c9fb) - no other reference to 1.1.1 remains in the repo outside historical reports - the ref-vs-pin guard at postprocessor.sh:164 still reads the same variable - `bash -n` clean on both files BOTH LAUNCHERS, deliberately. un_crafd carries the identical pin and the identical exposure; leaving CRAF'd on 1.1.1 would fix one partner delivery and silently leave the other on the version with the defect. NOT pinned to `development`, which would have been the lazy choice and is worse than a moving target here: the :164 guard compares the installed ref against the pin STRING, so a branch name satisfies the check designed to catch staleness while installing whatever landed most recently. It would pass the guard and still be unreproducible for a partner-visible delivery. Expect a behaviour change: a delivery that previously completed can now stop. 1.4.0 adds two refusal situations and no new exception types. A DeliveryNotFindableError names every object that does not resolve and states explicitly when nothing in the run is servable. Edited under an explicit lift of the standing "do not modify run.sh" rule, granted for this change. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01GG2cY2HaqmUMpwUyR5V82K * docs(#439): record WHY the pin moved to 1.4.0, in both launchers /review-diff on the pin bump, and the finding is one the value hid: the rationale blocks in both launchers document 1.1.0 and then 1.1.1 in detail, and stopped there. The value said 1.4.0 with no recorded reason. These blocks are not commentary. `tools/launcher/postprocessor.sh:26` says "EVERY STEP BELOW IS A SCAR. Read the comment before reordering anything." A future reader would have found a version two moves ahead of its own history and no way to learn what it carries. Which is the exact pattern the views-postprocessing session named against itself twice today — fixing the visible half of a finding and leaving the half that bites. The value is the visible half. The scar record is the half that explains it. un_fao now records the 2026-09-29 incident as the reason: the findability guard in 1.1.1 checked TWO artefacts out of the 110 a run uploads, and by returned file id rather than by name, so the first-ever FAO delivery uploaded 109 of 110, logged "Postprocessor Run Completed", fired a success alert and was refused by views-faoapi. It also records that the producer half of that defect is views-pipeline-core#552, shipped in 3.3.4 and reached from PyPI, while this pin is the other half and is reached from a git tag — which is precisely why "is the fix released?" had two different correct answers and this line was missed for hours. un_crafd records the same, framed by the shared-prefix argument its own block already makes twice: both launchers install into one conda prefix, so moving only the leg that failed downgrades the environment out from under the armed FAO delivery (C-139). CRAF'd runs the identical `_assert_delivery_is_findable`, so the defect is byte-identical there and has simply not fired on that leg yet — the same sentence the 1.1.1 note had to write about #268. Also replaced the verification recipe. The old one greps for `success is not True`, which verifies the 1.1.0-era C-79 fix and is still correct but says nothing about what 1.4.0 adds. The new one greps `DeliveryNotFindableError` in `delivery/findability.py` — verifying the build by what it REFUSES rather than by what it claims, which is the lesson of the whole incident. No behaviour change in this commit; the pin value is unchanged from the previous one. ruff clean, 8037 passed, `bash -n` clean on both files. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01GG2cY2HaqmUMpwUyR5V82K --------- Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
…at delivered, and the FAO chain is a script (#524) * fix(#516): pin xarray and pandas — the fresh env built successfully and WRONG #516 said a fresh un_fao environment is unbuildable. Measured on 2026-09-29, it is not: it builds, on a different pandas MAJOR, in silence. declared: views-datafactory>=1.9.0,<2.0.0 + numpy>=1.26.4,<2.0.0 resolved: numpy 1.26.4, xarray 2025.12.0, pandas 3.0.6 the pod that produced the first FAO delivery: xarray 2024.3.0, pandas 1.5.3 Neither xarray nor pandas was pinned anywhere. The working versions existed only as hand-pins on a pod that has since been destroyed, so a fresh pod silently picks different ones. A build failure is loud; this is not, which makes it the more dangerous of the two shapes the issue conflated. views-datafactory requires xarray outright and pandas only in an optional extra that is not installed, so pandas arrives THROUGH xarray. xarray is the sole carrier — datafactory's own pyproject comment says so — and pinning the carrier fixes it. The bound is measured, not guessed. xarray 2024.3.0 is the LAST release that accepts pandas 1.x; 2024.5.0 moved to pandas>=2.0 and 2024.9.0 to pandas>=2.1. A tidy-looking `<2025` cap would therefore have been WRONG: the cliff is inside the 2024 line, not at the year boundary. Verified by resolving the file, not by reading version numbers. Declared identically in both postprocessors because they share one prefix (C-116); pinning only one lets whichever runs last decide pandas for both, which is the same trap the existing numpy pin already documents one line above. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01GG2cY2HaqmUMpwUyR5V82K * fix(#517,#523): install the appwrite extra, prove it at preflight, add --rehearsal Three findings from one /falsify audit of "we are ready for a new 40-lesson run" (FALSIFIED). All three had one cause: every fix that made the 2026-09-29 run work was applied BY HAND to the pod, and the pod was destroyed. The pod was the artefact and nothing in git described it. #517 — the publish path was never installed. The string "appwrite" appeared nowhere in pod_run_model.sh, so `_build_datastore` raised at publish time, AFTER the full training run. That is exactly how the first FAO delivery attempt failed. The extra is now requested explicitly with no version, so views-hydranet's own range still decides which pipeline-core is installed, and the client is IMPORTED in preflight — the whole script is written to fail early, and the one dependency that failed late was the one not checked there. Verified by a real resolve: appwrite 13.6.1 alongside views-hydranet 0.1.2 and views-pipeline-core 3.3.4, with pandas 1.5.3 and numpy 1.26.4 intact. The same resolve confirms the C-151 toolz override is necessary rather than folklore — without it the resolver lands on toolz 0.11.2. #523 — a deliberate cheap run was reachable only by deleting the guard against an accidental one. The >=300 floor stays; `--rehearsal <lessons>` is the escape hatch. It takes the count on the command line, because main.py has no hyperparameter override and a flag without a count still forces an edit to a tracked config. It patches the POD's clone only, then re-imports to CONFIRM the patch took — a substitution that silently missed would give a 300-lesson run wearing a rehearsal label, or the reverse. The marking matters more than the permission. A rehearsal's parquets are structurally identical to a production run's and pass every check including the runner's own, so a REHEARSAL file is written beside them and MANIFEST gains `mode:` and `config: PATCHED after checkout` — the git sha alone no longer describes the run. Found while reviewing this change, not in the audit: section 2 does not re-clone when .git exists, so a production run on a pod that had rehearsed would read the LEFTOVER patch. The floor still refused it, but blamed the committed config and advised --rehearsal — a correct refusal for the wrong reason, sending the operator to edit the wrong file. It now detects the dirty config and names it. Deferred, and it is the weaker half: this runner MARKS a rehearsal but cannot REFUSE to publish one, because the publish step is not in this script. That guard belongs in views-pipeline-core. An xfail(strict) stub for it was written and DELETED — its regex spanned the whole file under re.S and XPASSed by accident, i.e. the exact class of guard-that-cannot-fire this audit exists to find. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01GG2cY2HaqmUMpwUyR5V82K * fix(#518): the guards were decorative — execute the script, and write the missing warning An independent /falsify guard-mode audit of 53bfba9 reverted EVERY fix in that commit, kept the comments, and got 22/22 green. 20 of 24 mutations survived. 10 of 14 guards were DECORATIVE, 4 WEAK, none held. The mechanism is the lesson. Every assertion read pod_run_model.sh as TEXT, and that script is unusually well commented — each fix carries a paragraph naming the incident and quoting the exact strings. So THE BETTER THE COMMENT, THE WEAKER THE GUARD: deleting the code left the comment, and the comment satisfied the assertion. Deleting the appwrite extra from the install line was invisible because the comment above it names views-pipeline-core[appwrite]. Worst survivor: `os.environ.get("REHEARSAL_LESSONS") or ""` -> `or "40"`. One word. Every production run then patches itself to 40 lessons, the floor is dead, the leftover-patch detector is dead, and because the SHELL variable stays empty the MANIFEST says `mode: production` and no REHEARSAL marker is written. A 40-lesson model labelled a production delivery, suite fully green. AND ONE GUARD CERTIFIED A SAFETY PROPERTY THAT DID NOT EXIST. It claimed the RunPod guide documents the /workspace chmod trap. The guide did not: no "world-readable", no "network filesystem", no C-154, no #518. It passed because `/workspace` is on line 104 and `chmod` on line 180, joined by `.*` under re.S — the identical defect the previous commit message boasted of having found and deleted elsewhere. That is worse than no guard, because it stopped the next reader looking. The warning is now written (#518, C-154): /workspace is a network filesystem where chmod 600 returns success and does nothing, leaving a credential at mode 666 with no error to notice. What changed in the tests: - the config-check program is EXTRACTED FROM ITS HEREDOC AND EXECUTED against fixture models in real git repos. That block holds all the rehearsal/production logic and was wholly unguarded. Includes an adversarial fixture whose literal is patchable but whose get_hp_config() returns 300 regardless, so a "verification" that compared target against itself is caught. - remaining text assertions read COMMENT-STRIPPED source, so a comment can never stand in for code. - requirements guards assert SPECIFIER SEMANTICS via packaging.SpecifierSet — is 1.5.3 admitted, is 3.0.6 refused — instead of pin presence. `pandas>=3.0` passed the old one, and marker-gated pins (`; python_version < "3.10"`, inert on 3.11) defeated the identical-pins check while keeping the captured text identical. - the roster guard CALLS get_hp_config() instead of matching the literal, which is what pod_run_model.sh itself insists on, citing #501 "the guard that was not one". pod_run_model.sh warned against this exact failure in three places and the guards did it anyway. That is the argument for the independence rule: the author's own model named the trap three times in one commit and stepped into it. 35 tests, up from 22. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01GG2cY2HaqmUMpwUyR5V82K * fix(#509): floor views-datafactory at 1.13.0, and guard the shell arg parsing Two gaps found by replaying the guard-mode audit's mutations against the rewritten guards — the replay is the point, since the previous version passed 22/22 with every fix reverted. #509 — the pod installed views-datafactory>=1.9.0. The credential-handling fixes landed in 1.13.0: before it the client could carry a netrc credential across a redirect to another host and embed it in error messages. This script only ever runs on hardware we do not own, carrying exactly that credential, and the guide already prescribes >=1.13.0 explicitly while the script did not. A resolver picks 1.13.0 anyway today, but "the resolver will probably do the right thing" is precisely the reasoning that put xarray 2025.12.0 and pandas 3.0.6 into a fresh environment (#516). Verified: the floor still resolves to 1.13.0 with hydranet 0.1.2, pipeline-core 3.3.4, appwrite 13.6.1, pandas 1.5.3. The shell argument parser was unguarded. Executing the config-check program does not cover it — the flag never reaches Python if the shell refuses it first. Deleting the `--rehearsal)` case arm removes H1's escape hatch entirely and left every guard green, because the string survives in the header comment and in USAGE. Now executed against the real script: the flag and its count must be CONSUMED (distinguished from "unknown option" by which error appears), a missing count must be refused, a non-integer must be refused, and — the control, without which the first assertion proves nothing — a genuinely unknown option must still be rejected. Argument parsing precedes `mkdir -p "$OUT"`, so these cases exit before touching the filesystem. MUTATION REPLAY — every surviving mutation from the audit is now caught: X7 `or ""` -> `or "40"` (production runs patch themselves) 3 failed X8 --rehearsal becomes a bare flag 3 failed X22 the --rehearsal case arm deleted 3 failed X1 extra dropped from the install line, comment kept 1 failed X2 both appwrite imports commented out, prose comment kept 1 failed X5 2026-09-29 install order restored + reassuring comment 1 failed X6 the floor prints instead of exiting 1 failed X9 leftover-patch detector flipped to the wrong branch 1 failed X10 patch verification becomes `if target != target` 1 failed X11 the marker-writing block deleted 1 failed X12 stale-marker clear moved to the end of the run 1 failed X13 MANIFEST mode branches swapped 1 failed X14 both files pinned to xarray 2025.12 / pandas 3.0 8 failed X15 upper bounds dropped 8 failed X17 pins marker-gated inert on 3.11 6 failed X18 get_hp_config() returns 40, the 300 literal untouched 1 failed X20b toolz override replaced by a comment 1 failed control: the guide's chmod warning deleted 1 failed X16 (`xarray == 2024.3.0`, a stricter and CORRECT pin) previously failed with a message claiming no pin existed; it now passes, so the guard no longer raises a false alarm against a legal PEP 508 spelling. Two honest notes. X18 was recorded SURVIVED in the audit but had not applied — the config returns a dict literal, so the mutation's sed matched nothing. Re-applied correctly, it is caught. And X19 (replacing the credential chmod with an unrelated one) now passes, because the /workspace warning it was meant to defeat genuinely exists; the control above proves the guard fires when that warning is removed. 40 tests, up from 35. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01GG2cY2HaqmUMpwUyR5V82K * fix(#509): actually apply the 1.13.0 floor the previous commit only tested for The previous commit shipped the guard and not the fix. The sequence: edit the install line, write the test, verify the test fires by reverting the line with sed, then `git checkout -- tools/podrun/pod_run_model.sh` to undo the sed — which reverted to the last COMMIT, discarding the uncommitted 1.13.0 edit along with the sed. The commit then captured the test plus a 1.9.0 install line. Caught by that test on the next full-suite run, which is the whole argument for writing guards that execute: a text-matching guard for this would have been satisfied by the comment above the line, which names 1.13.0 and explains why. Also worth recording, because it made the failure hard to see for two runs: this suite runs pytest-randomly, so ordering varies between runs. Two earlier runs reported 8058 passed/209 skipped and 7758 passed/470 skipped/3 failed with an unchanged tree, and the three failures were ordering interference in pre-existing tests, not in anything here. With `-p no:randomly` the counts are stable at 8076 passed/209 skipped/5 xfailed and the only failure was this one. Verify against a fixed order before concluding a change is responsible for a red suite. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01GG2cY2HaqmUMpwUyR5V82K * feat(#499): script the FAO delivery chain — it existed only as commands typed on a pod The fourth instance of the pattern the /falsify audit found three of, and the one that mattered most. On 2026-09-29 the whole FAO delivery — eight forecasting runs, the rusty_bucket pool, the publish, the un_fao postprocessor — was typed by hand on a rented pod. It worked, the pod was destroyed, and nothing recorded what had been done. The audit missed it because it audited the script that exists (pod_run_model.sh) instead of asking what the delivery needs. pod_run_model.sh runs `-r calibration -t -e` and its header is accurate that it uploads nothing: it is Track A, the research frames. Nothing in tools/ drove Track B at all — `grep -rln "prediction_store\|r forecasting" tools/` returns one unrelated file. This is not a stopgap. fimbulthul, which #499 names for every Track B step, has been unreachable since 2026-09-25, so a pod is the only hardware there is. pod_run_model.sh gains `--forecast`: `-r forecasting -t -f`, and it SKIPS sections 4 and 5. Those are calibration deliverables — a forecast has one origin, so the 13-parquet count check would refuse it, and what the FAO chain consumes is the pooled ensemble output. Kept in this script rather than a copy because everything before section 3 is identical and already exercised; a second copy would be a second place for C-151 to rot. pod_run_fao_delivery.sh drives the chain. Two details that are not obvious and that cost real time to establish: - `-sa/--saved` on the ensemble is REQUIRED. Without it rusty_bucket refetches instead of pooling the member forecasts just written, and the eight runs are wasted. - the two legs need DIFFERENT INTERPRETERS. This pod builds a uv venv; tools/launcher/postprocessor.sh uses `conda shell.bash hook` / `conda create --prefix`. A pod can satisfy one and not the other, and step 4 is last — so a missing conda kills the delivery after every GPU hour is spent. Preflight now checks for it. `--preflight` is the mode that protects the money. At 300 lessons step 1 alone is ~16 GPU hours, and everything it checks is knowable in seconds: checkout, venv, netrc, the three publish secrets, conda, that un_fao's REGION resolves to land_gaul rather than a disarmed value (C-110), GPU, disk. It accumulates and reports every problem rather than dying on the first, because a round trip per problem is billed by the second. The publish secrets are checked HERE and not at first publish because views-pipeline-core's PredictionStoreConfig claims to read them "once at startup and fail loud … preventing silent failures after hours of training" and does not — it is called from _build_datastore, after training (views-pipeline-core#557). Until that moves, this preflight is the only check that happens before the money is spent. A rehearsal is marked at every level, and the marker says the uncomfortable part out loud: nothing downstream refuses a marked rehearsal (#523), so its undertrained forecasts DO reach the FAO shelf and ARE servable. That is deliberate — it is the only way to test the chain — and it is why the final stage reads back what landed with `tools.liveness`, by name, rather than inferring success from an exit code. On 2026-09-29 the publish reported success having written nothing (C-155). 19 guards, written after the lesson that 10 of 14 of the previous set were decorative. The parser and `--preflight` are EXECUTED; the remaining text assertions read comment-stripped source and cover only ordering facts that cannot be executed without a GPU, a store and a partner bucket. PODRUN_ROOT exists so preflight is exercisable off a pod; it relocates the workspace and relaxes no check. Every preflight check was proven able to fire, including by shadowing nvidia-smi with a failing stub and stripping conda from PATH — a check that cannot fail is the defect this repository keeps finding. Two defects in this script were found by its own tests: an unchecked `mkdir -p "$OUT"` that let it run on into a broken tee off a pod, and a lock message that reported "another delivery is in progress" for a directory that simply was not writable. What these guards do NOT cover, stated rather than implied: nothing here runs a forecast, publishes, or reaches the FAO. The first real exercise is a rehearsal on a pod. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01GG2cY2HaqmUMpwUyR5V82K * docs(#499): Phase 4b — the FAO delivery chain, and how to read its outcome Phases 4 and 5 document Track A (calibration predictions for research). The FAO delivery had no operator documentation at all, because until now it had no script. Leads with --preflight, because the two things most likely to stop the delivery are invisible until the end: the three Appwrite publish secrets, and conda — the postprocessor launcher requires it while the pod builds a uv venv, so a pod can satisfy the training leg and not the delivery leg, and that failure lands after ~16 GPU hours. Records two readings an operator would otherwise get wrong: DeliveryNotFindableError is 1.4.0 WORKING (it verifies by refusing, and names what it checked), and a rehearsal's forecasts genuinely reach the FAO shelf and are servable rather than being contained. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01GG2cY2HaqmUMpwUyR5V82K * fix(#509): floor views-datafactory in the postprocessor legs too, not just the pod install Found re-running the clean-checkout simulation after the delivery script landed: I floored views-datafactory at 1.13.0 in tools/podrun/pod_run_model.sh and left both postprocessor requirement files declaring >=1.9.0. Half a fix, and the half I left out is the leg that runs LAST — so it would have been the one still holding a client without the credential fixes at the point the delivery reaches a partner. The reasoning that applied to the pod install applies here unchanged: before 1.13.0 the datafactory client could carry a netrc credential across a redirect to another host and embed it in error messages. The postprocessor runs on the same rented hardware and fetches the actuals it curates land -> land_gaul against, so it holds that credential too. Nothing about "this is the postprocessor" makes that safer. Declared identically in both files for the same C-116 reason as the pins above them: one shared envs/views-postprocessing prefix, so whichever postprocessor runs last decides the version for both, and flooring one is the same as flooring neither. The guard is widened rather than duplicated: it now iterates the pod install and both requirement files, and was verified to fail independently for each of the three when that source alone is reverted to >=1.9.0. A guard that passes because two of three sources are correct is how this got shipped half-done in the first place. Resolve unchanged: views-datafactory 1.13.0, xarray 2024.3.0, pandas 1.5.3, numpy 1.26.4. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01GG2cY2HaqmUMpwUyR5V82K * fix(#509): apply the postprocessor floor the last commit only tested for — again Second time in one session, same mechanism, and it is worth recording rather than quietly amending: the previous commit shipped the widened guard and not the fix, because I verified the guard by reverting each source with `git checkout -- <file>` while the fix was still UNCOMMITTED. The checkout reverted to the last commit and took the fix with it. My own ship-it procedure says mutation-verification runs AFTER the commit and before the push, for exactly this reason — "the tree is clean, so a mutation can be applied and reverted with git checkout --". I violated it twice in a row, once for the pod install line and once for both requirement files. The order is the control, not the care. Both files now declare views-datafactory>=1.13.0, with the reasoning restored: before 1.13.0 the client could carry a netrc credential across a redirect to another host and embed it in error messages, and the postprocessor runs on the same rented hardware holding the same credential. It is also the LAST leg — the one that reaches the partner. The guard is hardened while I am here: it now strips comments from the requirements files too, not only from the shell script. un_fao's xarray note quotes the old `views-datafactory>=1.9.0` line as the state it is explaining, so a guard reading raw text either trips on prose or has to be written carefully around it — and being careful around comments is precisely how the decorative guards got shipped in the first place. Verification order corrected: committed first, mutation replay follows. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01GG2cY2HaqmUMpwUyR5V82K * fix(#509): record the views-datafactory split as a deferral, and close the hole it opens The repo's own hygiene guard caught the last commit: flooring views-datafactory to >=1.13.0 in the two postprocessor files while 34 model requirements stay at >=1.9.0 makes one package declared two ways, which C-116 says is then decided by run order rather than intent. A good guard, and it fired on a real inconsistency I introduced. Not unified, deliberately, and the reasons are asymmetric: - the 34 are model requirements unrelated to the delivery legs. Raising them is #509's own scope, not a delivery PR's, and it would put 34 unreviewed one-line edits in a branch about reproducing a pod. - one of them is `models/violet_visitor/requirements.txt`. Another session owns that model's contents and this session is instructed not to edit them. That instruction is not mine to set aside because it would be convenient for a guard. So: DEFERRED_PACKAGES, which the guard offers explicitly, with the trigger named as the guard also requires — #509 raising the remaining 34, at which point the entry is DELETED rather than amended, because the divergence it describes will not exist. The divergence is also less dangerous than the rule's general case: the two postprocessors share one prefix and agree with each other, the models resolve elsewhere, and the resolver picks 1.13.0 for all of them today regardless. The floor only forbids something lower. AND THE PART WORTH THE EXTRA TEST. A DEFERRED_PACKAGES entry is wider than it looks: it also exempts the package from test_no_dependency_is_declared_without_an_upper_bound. So while this deferral stands, nothing in the repo would notice `<2.0.0` being dropped from a views-datafactory line — and that rule exists because an unbounded internal package installs the next breaking major on the following monthly run. Buying a floor by silently selling a ceiling is not a trade worth making quietly. Closed with test_every_datafactory_declaration_keeps_its_upper_bound, which is narrower than the rule it stands in for — it says nothing about floors, only that a ceiling exists wherever this package is named — and which outlives the deferral harmlessly. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01GG2cY2HaqmUMpwUyR5V82K * fix(#499): clear stale calibration output on the forecast leg, and agree on the workspace Two findings from reviewing this branch before merging it. Both are cases of a rule the scripts already state, not applied where it also holds. STALE CALIBRATION ARTEFACTS. The forecast leg writes no parquets, so it skipped sections 4 and 5 — but a previous calibration run on the same pod leaves parquet/ and draws/ in the SAME output directory. The forecast MANIFEST would then sit beside 13 parquets from a different run type, and the guide's rsync copies the directory, so they come home as this run's output. pod_run_model.sh already reasons about exactly this twice ("Clear a previous attempt first … if they happened to sum to 13 the count check would pass while the manifest covered two different training runs") for the case where the current run DOES produce them. The case where it produces none is the same hazard, and skipping a stage is not the same as clearing it. THE TWO SCRIPTS MUST AGREE ON $ROOT. pod_run_fao_delivery.sh honours PODRUN_ROOT so its preflight can be exercised off a pod; pod_run_model.sh had /workspace hardcoded. The delivery script reads $ROOT/deliver/<model>/STATUS to decide whether a model may be pooled, so a relocated workspace in one and not the other means reading a STATUS the other never wrote — and a missing STATUS is indistinguishable from a model that failed, which would refuse a healthy roster. Benign on a pod, where both are /workspace; latent, and cheap to remove. Both guarded. Written before the merge rather than after, which is the only time reviewing your own diff is worth anything. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01GG2cY2HaqmUMpwUyR5V82K --------- Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
…loor was below the one it delegates to (#527) * fix(#499): the FAO preflight's disk floor was below the floor it delegates to Found by a /falsify pass on "we are ready for a 40-lesson forecasting run reaching faoapi", and it is my defect from #524. pod_run_fao_delivery.sh demanded 60GB for EIGHT models plus the pooled ensemble, while pod_run_model.sh — which it delegates to eight times — refuses below 40GB for ONE model ("one model needs ~20GB"). So the orchestrator's preflight could report ready and the delegated script could then refuse partway through the roster, after hours of paid GPU time. That is precisely the failure --preflight exists to prevent, introduced into --preflight. Raised to DISK_FLOOR_GB=80: 40GB transient, which the delegated script enforces per model regardless of what is written here, plus eight models' retained output and the pool. The retained component is an ESTIMATE and the comment now says so rather than implying a measurement. One model's calibration output is ~2.5GB per predictions directory on this machine, but NO FORECASTING RUN HAS EVER COMPLETED ON THIS ROSTER — rusty_bucket's forecasting_log.txt is from 2026-07-20 and lists temporary_crane and temporary_fox at Deployment Status: shadow, a different roster entirely — so the retained size of a forecast is unmeasured. A forecast has one origin against calibration's thirteen so it should be smaller, and "should be" is doing real work in that sentence. Trigger for revising the number: the first completed run, rather than reasoning about it a second time. Guarded by a test that compares this floor against the delegated script's, so the two cannot drift apart again — the defect was not the number, it was that nothing related them. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01GG2cY2HaqmUMpwUyR5V82K * fix(#499): preflight checked 3 of 9 publish variables — build the config instead of counting Found by the /falsify pass on the forecasting claim, acting on a suggestion from the views-pipeline-core session. My preflight from #524 checked three environment variables and concluded the publish would work. Measured against the real code, PredictionStoreConfig requires NINE: APPWRITE_ENDPOINT, APPWRITE_DATASTORE_PROJECT_ID, APPWRITE_DATASTORE_API_KEY, APPWRITE_PROD_FORECASTS_BUCKET_ID, APPWRITE_PROD_FORECASTS_BUCKET_NAME, APPWRITE_PROD_FORECASTS_COLLECTION_ID, APPWRITE_PROD_FORECASTS_COLLECTION_NAME, APPWRITE_METADATA_DATABASE_ID, APPWRITE_METADATA_DATABASE_NAME Only the first three are secrets. A missing identifier fails the publish exactly as hard as a missing secret, so the check would have reported ready with six of nine absent — and the run would have failed at the publish, after the entire roster trained. The precise failure this mode exists to prevent, in the mode that exists to prevent it. Twice in one night now, which is its own signal. Adding the six missing names would not fix it. The extra can be absent, the endpoint unreachable, the key expired (#359: 2026-11-17). Preflight now CONSTRUCTS the config and imports the SDK, which answers the actual question instead of a proxy for it. views-pipeline-core suggested this and observed it is what #557 argues the manager should do itself; until it does, doing it by hand costs seconds. Second finding from the same measurement, and an operator trap: pipeline-core no longer auto-loads a .env from the working directory (#346, register C-177) — a library reading whatever .env its caller is standing in is what the Appwrite seam contract §3 forbids. So a correct /root/.secrets sitting beside the operator is NOT enough; the variables must be exported. Without that line the failure presents as wrong credentials rather than unexported ones. Preflight now prints `set -a; . /root/.secrets; set +a`. Both guarded, and the guards assert the construction and the exported-variables note exist rather than re-listing nine names — a guard that enumerates the same thing the code does is the defect that produced this commit. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01GG2cY2HaqmUMpwUyR5V82K --------- Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
This is the views-models release:
development→main, 442 commits, the first since 2026-06-26. Per #282 §1, a release of this repo is exactly this merge plus a tag — nothing is published, nothing installs it, and no sibling repo depends on it. Its artifact is the forecast, which is why the gate is a real run rather than an install check.Release gate — measured today, not inherited from the runbook
developmentmain ahead of development= 0; no partition numbers movedensembles/pink_ponyclub/run.sh -mgets past import on a real machinemonthly_run.shend to end, env snapshots committedunfao_bucket0.3.0, branch protection onmainG0 — the blocker is gone
The runbook recorded G0 at 4/5, blocked by "the one item nobody in this repo can fix" —
violet_visitor's committedloss_reg. #297 and #336 are both closed and all five workflows pass ondevelopment.G1 — re-verified, and one claim corrected
main ahead of developmentis 0. On partitions, a naive number-scrape of the 132 changedconfig_partitions.pyfiles reports 87 "moved", and that reading is wrong. Two things account for all of it:ViewsMonth.now().id(aningester3dependency) replaced by a local_current_month_id()using the January-1980 epoch — the1980is the epoch constant, not a boundary. Equivalence checked:(2026−1980)×12+8 = 560, matching the liveness surface'snow_month_id: 560.# ViewsMonth reference: 121 = Jan 1990, 444 = Dec 2016, ….The partition tuples themselves are byte-identical on both sides:
0 partition boundaries moved. The forecast windows this release ships are the ones
mainalready has.G4 — corrected
First reading, which was wrong.
python -m tools.livenessreports:That was taken at face value and this PR was opened claiming the release was blocked.
What is actually in the bucket, listed directly:
The forecast is there — 110 shards, delivered 2026-08-13. G4's delivery half is met.
The detector cannot see it.
tools/liveness/unfao_delivery.py:40-41:The delivery writes
rusty_bucket_forecasting_*, so the forecast half has never matched a real file, while the historical half matches and reports correctly. Those 110 shards are counted asother_files. Filed separately; it matters beyond this PR because #320 is "FAO forecast delivery has been stalled for 145 days and nothing detected it" — this detector exists for that, and is blind in the stream it was built to watch.What is still unverified is G4b: that faoapi can serve the delivered forecast. That is the half this repo cannot check alone.
Deletions — 2, both intended
Superseded by
tools/{scaffold,catalogs,partitions,audit}/under C-60. The runbook anticipated 9; 7 of those already flowed tomain.Still outstanding at G5, from the runbook
0.3.0— bareMAJOR.MINOR.PATCH, nov. Thevprefix is an active defect:git tag --list | sort -V | tail -1returnsv0.0.1as "newest" because the prefix sorts after digits.main, requiring the five workflows.maincurrently has none, which is how it took four direct hotfix commits in June 2026 that never flowed back.mainmeans. CLAUDE.md's own test applies: a rule a contributor could violate without knowing it existed.What this release explicitly does not do
Publish anything · restructure the 131-requirements mismatch (C-116). Two items that were on this list are done and no longer gate anything here: the views-r2darts2 three-way split (C-115) was closed by #461 — all 31 declare one spec; and both views-hydranet (0.1.0, 2026-09-15, views-hydranet#340; the 8 models pin
~=0.1.0, #462) and views-postprocessing (1.2.0) are on PyPI. The FAO launcher still installs views-postprocessing fromgit+…@main— a moving branch pointer, tracked with views-pipeline-core#222 — but PyPI availability is no longer the reason. (Updated 2026-09-15, #468.)