feat(#505): q95 and E[y|y>0] as selector frames — and the separation that keeps them out of the contract - #521
Open
Polichinel wants to merge 1 commit into
Open
Polichinel wants to merge 1 commit into
Polichinel wants to merge 1 commit into
Conversation
…that keeps them out of the contract The #505 experiment needs one posterior cube collapsed three ways and submitted as three apparent models, so the ensemble selector's own criteria choose between a mean that under-predicts total fatalities ~5x and two estimators that do not. The converter had only `arithmetic_mean` and `median`. Adds `q95` and `conditional_mean` (`E[y|y>0]`, written as `sum / n_positive` — the form of `E[y] / P(y>0)` that cannot divide by a zero probability). On the 2026-09-28 calibration run both come out ~4x the mean across all 7 landed models, which is the right order to close that gap. **The design point is what did NOT change.** `AGGREGATE_METHODS` stays exactly the two names views-hydranet ADR-021 defines, because that tuple is contract vocabulary — the only values a model may *declare*, enforced against every roster member by test_collapse_declaration_matches_the_converter. The new estimators go in a separate `ESTIMATORS` registry, a strict superset. Merging the two would be the whole defect: a config could declare `q95` and the roster test whose job is to catch that would wave it through as a legitimate ADR-021 method. Pinned by a test asserting strict subset and naming the difference. `aggregate_method` -> `estimator` on the API, since for two of four values the old name was simply false. `--aggregate-method` stays accepted so runbooks keep working. Acceptance condition, met: the mean frames this code produces are byte-exact against the parquets delivered before it — 7 models x 13 origins, pd.testing.assert_frame_equal with check_exact. The refactor moved nothing on the delivered path. 8 mutations applied and all caught: dividing by D instead of n_positive; NaN instead of 0.0 for an all-silent cell; counting zeros as positive; q95 as the median; q95 as the max; declaring q95 an aggregate method; unknown estimator falling back to the mean; silently reordering rows. Tree restored byte-identical after each. ADR-023 Amendment 1 records the standing rule (a new estimator is a key in `ESTIMATORS` and nothing else; `AGGREGATE_METHODS` moves only when ADR-021 does), that provenance lives in the output directory and nowhere else, and that q95 at D x K = 16 is a coarse tail estimate reported rather than corrected for. 289 passed across collapse, roster and delivery-coherence suites; ruff clean. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01GG2cY2HaqmUMpwUyR5V82K
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
The #505 experiment needs one posterior cube collapsed three ways and submitted as three apparent models, so the ensemble selector's own criteria choose between a mean that under-predicts total fatalities ~5× and two estimators that do not. The converter had only
arithmetic_meanandmedian.This code produced the 24 frames delivered to researchers on 2026-09-29. It has been sitting unpushed on one laptop since; the parquets are safe, the tool that made them was not.
What it adds
q95conditional_meanE[y|y>0]— mean over the positive draws, written assum / n_positive: the form ofE[y] / P(y>0)that cannot divide by a zero probabilityMeasured on the 2026-09-28 calibration run, across all 8 models: q95 is 4.55× the mean, conditional_mean 4.68×, against a measured ~5× under-prediction. Which is the point of the experiment.
The design point is what did NOT change
AGGREGATE_METHODSstays exactly the two names views-hydranet ADR-021 defines, because that tuple is contract vocabulary — the only values a model may declare, enforced against every roster member bytest_collapse_declaration_matches_the_converter.The new estimators go in a separate
ESTIMATORSregistry, a strict superset.Merging the two would be the whole defect: a config could declare
q95, and the roster test whose job is to catch exactly that would wave it through as a legitimate ADR-021 method. Pinned by a test asserting strict subset and naming the difference.The API parameter is renamed
aggregate_method→estimator, since for two of four values the old name was simply false.--aggregate-methodstays accepted so runbooks keep working.Acceptance condition, met
The mean frames this code produces are byte-exact against the parquets delivered before it — 8 models × 13 origins,
pd.testing.assert_frame_equal(check_exact=True). The refactor moved nothing on the delivered path.Mutations — 8 applied, 8 caught, tree byte-identical after each
n_positiveq95an ADR-021 aggregate methodDocumentation
ADR-023 Amendment 1 records the standing rule — a new estimator is a key in
ESTIMATORSand nothing else;AGGREGATE_METHODSmoves only when ADR-021 does — plus that provenance lives in the output directory and nowhere else, and that q95 at D×K=16 is a coarse tail estimate reported rather than corrected for.Verification
289 tests pass across the collapse, roster and delivery-coherence suites.
ruffclean.The delivered output was separately audited in full — 312 parquet, 728,035,776 rows, no NaN/inf/negative, no duplicate keys, every file a complete 64,818 × 36 block, all three estimators live on the same 1,469,004 cell-months. Maps inspected and conflict-shaped.
🤖 Generated with Claude Code
https://claude.ai/code/session_01GG2cY2HaqmUMpwUyR5V82K