Phase 0, task 0.7 — see docs/studio/plan/PHASE_0.md.
The first task is a MEASUREMENT, not an assertion. Re-run two archived cases at today's submodule SHAs and record the observed per-quantity deviation in studio/tests/golden/REFERENCE_TOLERANCES.md, together with those SHAs. Bit-for-bit reproduction is a hypothesis (ASSUMPTION-2), not a premise: the archived ensemble has no provenance record and the solver has changed since it was produced.
- Tier A (
-m tier_a, CI, seconds): schema/units/DAG/hash/expansion tests + one 1-day, 40-bin run.
- Tier B (
-m tier_b, nightly/manual): 4–6 curated cases across D1/D2/D3/burst × sabr220/sabr330.
Fixtures store a documented uniform-stride reduction, not the 3.6 MB npz — uniform because an adaptive grid aliases the nucleation burst by up to 8×.
Do not re-baseline a deviation into a passing test without recording what it was and why.
Base branch: studio/dev.
Phase 0, task 0.7 — see
docs/studio/plan/PHASE_0.md.The first task is a MEASUREMENT, not an assertion. Re-run two archived cases at today's submodule SHAs and record the observed per-quantity deviation in
studio/tests/golden/REFERENCE_TOLERANCES.md, together with those SHAs. Bit-for-bit reproduction is a hypothesis (ASSUMPTION-2), not a premise: the archived ensemble has no provenance record and the solver has changed since it was produced.-m tier_a, CI, seconds): schema/units/DAG/hash/expansion tests + one 1-day, 40-bin run.-m tier_b, nightly/manual): 4–6 curated cases across D1/D2/D3/burst × sabr220/sabr330.Fixtures store a documented uniform-stride reduction, not the 3.6 MB npz — uniform because an adaptive grid aliases the nucleation burst by up to 8×.
Do not re-baseline a deviation into a passing test without recording what it was and why.
Base branch:
studio/dev.