From f42deae6a86e0dd7f2d581f1ad326fb63cd60d65 Mon Sep 17 00:00:00 2001 From: Polichinl Date: Tue, 29 Sep 2026 22:41:44 +0200 Subject: [PATCH] =?UTF-8?q?docs(falsify):=2040-lesson=20run=20readiness=20?= =?UTF-8?q?=E2=80=94=20FALSIFIED,=20two=20hard,=20two=20soft?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Claim audited: "we are ready for a new full (40-lesson) run." HARD 1 — views-models still pins VIEWS_POSTPROCESSING_PIN="1.1.1" in both launchers. A run launched now installs the build whose findability guard checks 2 of 110 artefacts and cannot see the sidecar by construction — the exact defect that made today's delivery unservable. 1.4.0 exists, is tagged, is on PyPI, and is referenced by nothing. views-models#439 carries the one-line diff. HARD 2 — a fresh launcher environment is unbuildable (views-models#516, open). xarray is unpinned in postprocessors/{un_fao,un_crafd}/requirements.txt, and views-datafactory's `xarray>=2024.1,<2026` cap permits 2025.12.0, which requires pandas>=2.1 against the platform's pandas 1.5.3. Today's pod works only because xarray was hand-pinned to 2024.3.0 on it. So the claim holds only if the run reuses that exact pod and its env survives the launcher's dry-run check. SOFT 3 — the SELECTION half of the findability guard still depends on an ordering guarantee that does not exist. #314 made the per-object half order-independent; the legs loop still resolves through pipeline-core's get_latest_file_id, which documents "newest by creation timestamp" and takes files_list[0] from an unsorted result. Held since August, so order has been favourable rather than guaranteed, and the set it indexes into grows ~120 documents per run at 40 lessons. The failure mode is a FALSE invisible-delivery on a healthy run. Stub S7 below. SOFT 4 — ADR-013 §4.6's capacity inequality (assembled-run size x safety factor <= consumer serving RAM) is an open maintainer item, and its reference figure is 28.6 GB at 36 months, already stated as larger than the serving host's RAM. 40 lessons raises it ~11%. Whether that matters depends on the run's S, which this seat does not know — recorded rather than asserted. PASSED — P4, the one I most expected to break: the new per-object query works on the live path. `_build_partner_read_store` sets store.model_path = None, and get_predictions_by_metadata injects `filters["name"]` only when model_path has a model_name, so the C-77 suppression holds for documents() as well as for latest_file_id. Verified in pipeline-core at datastore.py:423-432. Only S7 is assertable in this repo. HARD 1 and HARD 2 are other repositories' state and views-models is not in this repo's CI sibling checkout (ADR-016), so a test reaching for them would pass vacuously — a guard that cannot fire (C-102). They are recorded in the module docstring with their issue numbers instead. Co-Authored-By: Claude Opus 5 Claude-Session: https://claude.ai/code/session_01ANY1CCy9Xo7zjMY4XJ69v9 --- tests/test_falsification_release_readiness.py | 54 +++++++++++++++++++ 1 file changed, 54 insertions(+) diff --git a/tests/test_falsification_release_readiness.py b/tests/test_falsification_release_readiness.py index ac4a2fc..4705c6b 100644 --- a/tests/test_falsification_release_readiness.py +++ b/tests/test_falsification_release_readiness.py @@ -146,3 +146,57 @@ def test_the_publish_job_cannot_ship_untested_code(): # and are recorded in the sprint rather than as assertions: views-faoapi is blocked # downstream (#294), and #272's second question is unanswered. # ───────────────────────────────────────────────────────────────────────────── + + +# ───────────────────────────────────────────────────────────────────────────── +# Fourth audit, 2026-09-29. Claim: *"we are ready for a new full (40-lesson) run."* +# +# Verdict: FALSIFIED — two hard, two soft. Only ONE is assertable in this repo and +# it is below. The others are facts about other repositories' state, and a test +# here that reached for them would be a guard that cannot fire in CI (C-102): +# +# HARD 1 — views-models still pins VIEWS_POSTPROCESSING_PIN="1.1.1", so a run +# launched now installs the build whose guard checks 2 of 110 artefacts. +# Lives on views-models#439. Not assertable here: views-models is not in +# this repo's CI sibling checkout (ADR-016), so the assertion would pass +# vacuously. This is C-112's whole subject. +# HARD 2 — a fresh launcher environment is unbuildable (views-models#516, open): +# xarray is unpinned in the launcher requirements and datafactory's +# `>=2024.1,<2026` cap permits 2025.12.0, which needs pandas>=2.1 against +# the platform's pandas 1.5.3. Not ours and not assertable here. +# SOFT 4 — ADR-013 §4.6's capacity inequality is an open maintainer item and +# 40 lessons raises the assembled-run figure ~11%. Needs the run's S, +# which this seat does not know. Recorded, not tested. +# ───────────────────────────────────────────────────────────────────────────── + + +@pytest.mark.xfail(reason="S7: unaddressed falsification — see the audit report", strict=True) +def test_the_selection_check_does_not_depend_on_an_order_nothing_guarantees(): + """S7 (soft). The per-object half stopped depending on document order in #314. + The SELECTION half still does, through a guarantee that does not exist. + + `verify`'s legs loop calls the port's `latest_file_id`, which is pipeline-core's + `get_latest_file_id`. That function DOCUMENTS *"the file ID of the newest matching + file based on creation timestamp"* and implements `files_list[0]` over the result of + `search_files_by_metadata`, which appends only `Query.equal` per filter and never an + `order_desc`/`order_asc` (`modules/appwrite/file.py:1045-1050`). + + So the check that decides whether the consumer's selection lands on THIS run takes + an arbitrary element of every `category="forecast"` document ever written, and + compares it against this run's manifest id. It has held since August, which means + the order has been favourable rather than guaranteed — and the set it indexes into + grows by ~110 documents per run, ~120 at 40 lessons. + + The failure is a **false** `DeliveryNotFindableError` on a healthy delivery, which + is the direction this module exists to avoid. + + Fails until either the selection half is made order-independent here, or + pipeline-core's sort lands and the pin moves. Filed upstream after the #312 + re-review; the fix is not ours and the dependency is. + """ + source = (_REPO / "views_postprocessing" / "delivery" / "findability.py").read_text() + legs_loop = source[source.index("for category, expected in legs.items():"):] + assert "resolve_latest(" not in legs_loop.split("assert_all_findable")[0], ( + "the selection check still resolves through `latest_file_id`, i.e. through " + "pipeline-core's unsorted `files_list[0]`" + )