Skip to content

feat(#505): q95 and E[y|y>0] as selector frames — and the separation that keeps them out of the contract - #521

Open
Polichinel wants to merge 1 commit into
developmentfrom
feat/505-selector-frame-estimators
Open

Polichinel wants to merge 1 commit into
developmentfrom
feat/505-selector-frame-estimators

Conversation

@Polichinel

Copy link
Copy Markdown
Collaborator

The #505 experiment needs one posterior cube collapsed three ways and submitted as three apparent models, so the ensemble selector's own criteria choose between a mean that under-predicts total fatalities ~5× and two estimators that do not. The converter had only arithmetic_mean and median.

This code produced the 24 frames delivered to researchers on 2026-09-29. It has been sitting unpushed on one laptop since; the parquets are safe, the tool that made them was not.

What it adds

estimator definition
q95 95th percentile across the 16 draws, numpy's linear interpolation
conditional_mean E[y|y>0] — mean over the positive draws, written as sum / n_positive: the form of E[y] / P(y>0) that cannot divide by a zero probability

Measured on the 2026-09-28 calibration run, across all 8 models: q95 is 4.55× the mean, conditional_mean 4.68×, against a measured ~5× under-prediction. Which is the point of the experiment.

The design point is what did NOT change

AGGREGATE_METHODS stays exactly the two names views-hydranet ADR-021 defines, because that tuple is contract vocabulary — the only values a model may declare, enforced against every roster member by test_collapse_declaration_matches_the_converter.

The new estimators go in a separate ESTIMATORS registry, a strict superset.

Merging the two would be the whole defect: a config could declare q95, and the roster test whose job is to catch exactly that would wave it through as a legitimate ADR-021 method. Pinned by a test asserting strict subset and naming the difference.

The API parameter is renamed aggregate_method → estimator, since for two of four values the old name was simply false. --aggregate-method stays accepted so runbooks keep working.

Acceptance condition, met

The mean frames this code produces are byte-exact against the parquets delivered before it — 8 models × 13 origins, pd.testing.assert_frame_equal(check_exact=True). The refactor moved nothing on the delivered path.

Mutations — 8 applied, 8 caught, tree byte-identical after each

mutation
divide by D instead of n_positive caught
NaN instead of 0.0 for an all-silent cell caught
count zeros as positive caught
q95 as the median caught
q95 as the max caught
declare q95 an ADR-021 aggregate method caught
unknown estimator falls back to the mean caught
silently reorder rows caught

Documentation

ADR-023 Amendment 1 records the standing rule — a new estimator is a key in ESTIMATORS and nothing else; AGGREGATE_METHODS moves only when ADR-021 does — plus that provenance lives in the output directory and nowhere else, and that q95 at D×K=16 is a coarse tail estimate reported rather than corrected for.

Verification

289 tests pass across the collapse, roster and delivery-coherence suites. ruff clean.

The delivered output was separately audited in full — 312 parquet, 728,035,776 rows, no NaN/inf/negative, no duplicate keys, every file a complete 64,818 × 36 block, all three estimators live on the same 1,469,004 cell-months. Maps inspected and conflict-shaped.

🤖 Generated with Claude Code

https://claude.ai/code/session_01GG2cY2HaqmUMpwUyR5V82K

…that keeps them out of the contract

The #505 experiment needs one posterior cube collapsed three ways and submitted as
three apparent models, so the ensemble selector's own criteria choose between a mean
that under-predicts total fatalities ~5x and two estimators that do not. The converter
had only `arithmetic_mean` and `median`.

Adds `q95` and `conditional_mean` (`E[y|y>0]`, written as `sum / n_positive` — the form
of `E[y] / P(y>0)` that cannot divide by a zero probability). On the 2026-09-28
calibration run both come out ~4x the mean across all 7 landed models, which is the
right order to close that gap.

**The design point is what did NOT change.** `AGGREGATE_METHODS` stays exactly the two
names views-hydranet ADR-021 defines, because that tuple is contract vocabulary — the
only values a model may *declare*, enforced against every roster member by
test_collapse_declaration_matches_the_converter. The new estimators go in a separate
`ESTIMATORS` registry, a strict superset. Merging the two would be the whole defect: a
config could declare `q95` and the roster test whose job is to catch that would wave it
through as a legitimate ADR-021 method. Pinned by a test asserting strict subset and
naming the difference.

`aggregate_method` -> `estimator` on the API, since for two of four values the old name
was simply false. `--aggregate-method` stays accepted so runbooks keep working.

Acceptance condition, met: the mean frames this code produces are byte-exact against
the parquets delivered before it — 7 models x 13 origins, pd.testing.assert_frame_equal
with check_exact. The refactor moved nothing on the delivered path.

8 mutations applied and all caught: dividing by D instead of n_positive; NaN instead of
0.0 for an all-silent cell; counting zeros as positive; q95 as the median; q95 as the
max; declaring q95 an aggregate method; unknown estimator falling back to the mean;
silently reordering rows. Tree restored byte-identical after each.

ADR-023 Amendment 1 records the standing rule (a new estimator is a key in `ESTIMATORS`
and nothing else; `AGGREGATE_METHODS` moves only when ADR-021 does), that provenance
lives in the output directory and nowhere else, and that q95 at D x K = 16 is a coarse
tail estimate reported rather than corrected for.

289 passed across collapse, roster and delivery-coherence suites; ruff clean.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GG2cY2HaqmUMpwUyR5V82K
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant