fix(heldout): declare score-reliability sample size - #1100
Conversation
Reliability evidence must fail closed without a positive sample_size, and the declared count must be the person population per case.
Remove the hidden 1,200-row reliability default. The harness run still writes 1,200 as this run's choice and records sample_size_per_case.
ADR 0051 is Proposed. Production route/conduct defaults stay locked.
|
Important Draft PR not reviewedDraft PRs are not automatically reviewed by default.
To automatically review draft PRs, update your CodeRabbit configuration: reviews:
auto_review:
drafts: trueThanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
Signed-off-by: Seongho Bae <me@seonghobae.me>
|
Normal predecessor integration at 3e85d32 retains original reliability sample declaration 9ff12b5 and includes predecessor da4ceac. No-detection censoring, alarm/detection denominators, null ineligible calibration candidate and retained calibration results now propagate through this stack. Merged-tree validation: 31 boundary/ADR identifier tests passed in 12.92 seconds; Ruff and diff whitespace checks passed; CodeGraph indexing completed. No force push or discarded delta. This is local regression evidence, not real-data accuracy, protected merge or release evidence. Required checks and reviews must use the new head. |
Stack
Draft successor of #1099 (
b975c3dd). Parents #1098, #1097, #1096, #1095, #1093, #1092, #1091, #1090, #1074, and psychometric KPI owner #1067 stay open. This PR does not close #1099.Exact head:
9ff12b5dfbf2a36f21f0e16cc311eb3949b90e94.Problem
Held-out posterior reliability evidence hid
RELIABILITY_SAMPLE_SIZE = 1_200, so weak- versus strong-information cases used a repository-selected person population.Change
sample_sizeon_validate_score_reliability. Missing, boolean, or non-positive values fail closed.sample_size_per_case.Verification
Local smallest tests:
tests/test_psychometric_routing.py,tests/test_psychometric_benchmark_boundaries.py,tests/test_planning_adr_identifiers.py)This is declaration-contract evidence, not buyer-held-out accuracy, p95 latency, or protected merge.
Remaining no-heuristics work
Other harness sample sizes (judge, exploration rate) remain later slices. No production policy or release is authorized. Parent #1067 still needs independent review.