Skip to content

feat(routing): make psychometric evidence fail closed - #1064

Open
seonghobae wants to merge 149 commits into
ContextualWisdomLab:mainfrom
seonghobae:codex/psychometric-uncertainty-20260905
Open

feat(routing): make psychometric evidence fail closed#1064
seonghobae wants to merge 149 commits into
ContextualWisdomLab:mainfrom
seonghobae:codex/psychometric-uncertainty-20260905

Conversation

@seonghobae

@seonghobae seonghobae commented Sep 4, 2026

Copy link
Copy Markdown
Contributor

Successor tracking

The complete delta at this PR's head ba010d5a023f69204df73e0a8e4d9b0a5bf7ee2f is preserved in trusted-branch successor #1067 at 1481c595dc1d16e7bf4b65addaf0bd30322cf2b8. Commit ancestry and tree equality are recorded in docs/product-technical-gap-baseline.md. Review and required-workflow integration continue there. This PR remains open until protected delivery and a fresh full-delta audit; later changes here must also be carried forward.

Summary

  • distinguish measured failure from buyer evidence that has not run
  • enumerate scale-linking, recalibration- and candidate-roster-invariance, response-pattern, local-independence, DIF, item-side, judge, uncertainty, and adaptive-exposure validity requirements
  • retain secret-free deterministic selection-design receipts without inventing propensity
  • validate probability calibration, positive-propensity logging, common-item linking, parameter, functional, and sequential drift, alternate-form score equating, candidate DIF, judge severity, an item-side covariate, parameter uncertainty, construct dimensionality, global model fit, score reliability, cross-condition generalizability, conditional information, classification decisions, decision utility, and distinct predictive-fit axes in the synthetic benchmark
  • classify IRT-Router as an IRT-shaped predictor rather than validated invariant measurement
  • keep the production routing gate closed until buyer-held-out accuracy, latency, and measurement validity pass

Exact-head evidence

  • head: ba010d5a023f69204df73e0a8e4d9b0a5bf7ee2f
  • focused routing/reliability/streaming suite: 200 passed on predecessor head
  • predecessor psychometric and paper contract suite: 32 passed in 15.50s
  • exact-head full repository suite: 3417 passed, 2 skipped in 666.99s
  • git diff --check: pass

Measurement boundary

The local synthetic benchmark preserves its accuracy improvement, but its latency confidence interval does not authorize production. Live deterministic routing records candidate, attempt, selection, and policy identity with propensity_status=not_identified; no counterfactual correction is claimed.

A fixed-seed benchmark-only 20% epsilon-greedy design assigns all four candidates probability at least 0.05 over 24,000 trials. Naively averaging the adaptively selected observations gives RMSE 0.321979 against known truth; logged inverse-propensity estimation reduces RMSE to 0.008943, an improvement of 0.313036, and all four 95% intervals cover known truth. Six-anchor Stocking-Lord linking recovers slope 1.3 and intercept -0.4 with true-parameter RMSE 3.24e-16. Purified logistic DIF detects the one injected cohort shift with recall 1.0 and zero false positives. A connected many-facet Rasch fit recovers three judge severities with RMSE 0.018292.

Two-neighbor interpolation reduces held-out logit calibration RMSE from 0.231235 to 0.025803; the paired candidate-minus-baseline 95% interval is [-0.208030, -0.202924]. Calibration slope moves from 0.991445 to 1.014030. These known synthetic probabilities do not establish calibration on buyer outcomes.

The released item-side covariate fit converges after 941 iterations and estimates -0.789650 for true delta=-0.8 (absolute error 0.010350). The released Oakes-information API gives item-intercept RMSE 0.039160, 95% interval coverage 1.0, and mean interval width 0.295945 for six known intercepts. It conditions on population parameters and rejects anchors, zero inflation, and item covariates. A separate recalibration-invariance screen links through seven stable anchors and flags the one injected drift item with no stable-item false positive; its fixed 0.25 tolerance is an effect-size rule, not a significance test. Nonparametric person fit ranks one injected inverted candidate pattern first among 1,000 patterns with ZU3 separation 1.818719, but does not infer a cause or universal cutoff. Horn parallel analysis retains both known dimensions in 1,000 binary response vectors across 12 items; because Pearson-PCA screening is not construct identification or confirmatory fit, buyer dimensionality remains unexecuted. Limited-information M2 accepts the fitted one-factor synthetic case at p=0.105619 and detects known two-factor misspecification at p≈2.27e-41; sensitivity varies by misspecification, so global buyer model fit stays unexecuted. Posterior empirical reliability rises from 0.366437 to 0.800436 when true item discrimination rises from 0.45 to 1.5; reliability remains separate from model fit and validity. Spreading 12 item difficulties across [-2, 2] improves worst information over trait points [-2, 0, 2] by 15.789% and lowers worst conditional standard error from 0.890897 to 0.827931, while exposing the center-precision tradeoff. At a declared synthetic cut of zero, lowering score standard error from 0.8 to 0.2 raises expected classification accuracy from 0.814182 to 0.996895 and consistency from 0.710275 to 0.993829; no buyer cut or asymmetric decision cost is inferred. These synthetic checks validate calculation contracts only; they do not establish buyer outcomes, live randomization, invariant scales, buyer uncertainty, construct validity, or production authorization.

A candidate-roster screen independently calibrates 20- and 16-candidate synthetic rosters and links 200 common items. The retained candidates have linked-score RMSE 0.010866, correlation 0.999999, and maximum shift 0.016498. Versioned buyer rosters and preregistered shift targets remain absent, so this does not establish sample-free buyer measurement.

A two-facet G-study separates candidate, query, occasion, and interaction variance in an 80-candidate, 12-query, four-occasion synthetic tensor. The D-study raises dependability from 0.401565 for one query and one occasion to 0.730184 for six and two and 0.849616 for 12 and four. A complete balanced buyer design, random-facet justification, and a registered target remain absent.

A decision-utility screen raises synthetic predictive validity from 0.2 to 0.6, increasing Taylor-Russell selected success from 0.500273 to 0.723515 and Brogden-Cronbach-Gleser net utility from 2,042.21 to 5,626.64. Holding validity fixed while increasing total measurement cost from 2,000 to 10,000 leaves selected success unchanged but makes net utility negative at -2,373.36. This personnel-selection model is an analogue, not validated routing economics; buyer-valued outcomes, actual costs and volume, selection ratio, and a preregistered utility target remain absent.

The predictive-fit report now distinguishes held-out queries for known candidate deployments from a held-out candidate deployment. Existing synthetic Brier and log-loss results belong only to the first task. Executing the unseen-candidate path across 24 contexts produces zero psychometric prediction coverage because the router does not fabricate an unobserved candidate score. This measured failure prevents query holdout evidence from being presented as cold-start deployment generalization.

A separate few-observation onboarding screen reuses released Rust-backed maximum-information EAP selection across 400 known synthetic candidates and a 31-query bank. It reaches target SE 0.5 after 7.1775 queries on average versus 10.47 for a seeded random order. Paired 95% intervals are [-3.4125, -3.18] queries, [-0.080874, 0.006129] theta squared error, and [-0.008804, -0.005716] unobserved-probability squared error. The theta interval includes zero, so no general theta-accuracy gain is claimed. The maximum 12 calibration calls measure onboarding burden, not live routing latency; zero-observation coverage and production authorization remain unchanged.

A classification-oriented screen stops when a 95% normal interval excludes a declared zero cut, with the same 12-query maximum. Across 400 known synthetic candidates, it stops early for 41%, averages 9.875 queries, and exactly matches the fixed-length decisions and 0.9125 accuracy. Paired intervals are [-2.425, -1.835] queries and [0, 0] accuracy. Buyer cuts and costs, near-cut risk, interval calibration, and live provider latency remain absent, so this does not authorize production.

Distance-from-cut strata expose the aggregate limit. Within 0.5 of the synthetic cut, candidates stop early only 3%, average 11.86 queries, and reach 0.70 accuracy. Candidates at least 1.0 away stop early 68%, average 8.305 queries, and reach 1.0 accuracy. These descriptive known-truth strata are not calibrated buyer subgroup guarantees.

Keeping unresolved cases explicit changes the interpretation: 42.5% of candidates are confidence-resolved and all resolved synthetic decisions are correct, while only 3% of near-cut candidates resolve. The remaining cases stay unresolved rather than being treated as production-ready decisions; buyer fallback policy and calibrated coverage remain absent.

A reject-option frontier now separates threshold selection from evaluation. Maximizing development coverage subject to a 2.5% Wilson 95% selective-risk upper bound selects z=1.645. Against z=1.96 on the same independent responses, coverage rises from 44.25% to 56% and all-candidate mean queries fall from 9.88 to 8.395. Paired 95% intervals are [8.75, 14.75] percentage points and [-1.715, -1.2625] queries. Observed selective risk is zero with a 1.686% Wilson upper bound, and directional coverage differs by 3 points. A ten-seed replication audit rejects this threshold despite the single-run efficiency. Coverage gain remains 9.75–15 percentage points and all-candidate query reduction remains 1.27–1.485, but selective risk reaches 1.802% and its Wilson upper bound reaches 4.540%. Only 20% of replications satisfy the declared 2.5% ceiling, so z=1.645 is marked rejected_not_replication_stable and cannot become a production threshold. The ten-run pass-rate MCSE is 0.1265; coverage-delta, all-candidate query-delta, and selective-risk MCSEs are 0.00619, 0.02257, and 0.00211. A conservative binomial design requires 400 replications for target pass-rate MCSE 0.025, so ten runs are a falsification screen rather than a precise operating-characteristic estimate.

A stdlib-only hot-path change replaces generator iteration in cosine dot products with map(operator.mul, ...), preserving math.fsum and multiplication order. Across ten fresh-process before/after runs, candidate p50 latency falls from a median 0.01325 ms to 0.007708 ms (41.83%), and the paired candidate-minus-baseline latency-CI upper bound falls from a median 0.000910 ms to 0.000635 ms (30.14%). The bound remains positive, so the decision-latency gate stays failed. These local CPU timings do not establish end-to-end provider latency.

A functional-impact screen also identifies the one item producing TCC-area difference 0.123355 and reduces the residual difference to zero. Its backward-only fixed-threshold search is narrower than Guo, Zheng, and Chang's complete stepwise method and does not supply a buyer cutoff.

A one-stream Bernoulli CUSUM screen measures the false-alarm and temporal-detection tradeoff separately. A bounded search over 11 thresholds uses 500 calibration runs and a 95% Wilson false-alarm upper bound, then evaluates selected threshold 6.6 on an independent 500-run seed. Held-out false alarms are 2.4% with upper bound 4.15%; detection-delay p50 is 10 and p95 is 20. These are synthetic calculation contracts, not Chen, Lee, and Li's multistream Bayesian compound-risk procedure or buyer-approved thresholds.

Alternate-form linear equating recovers a known slope 2 and intercept 1, reducing raw cross-form RMSE from 6.782330 to zero. Three hundred bootstrap 95% intervals cover all 11 known equivalent scores, but equal synthetic form populations do not establish buyer score comparability.

Owner dependency

Local-independence indices remain proposed in ContextualWisdomLab/fast-mlsirm#1748 at exact head 8461914a5bf04f9732add77761dd121bbec00103. A detached exact-head audit passes 44 focused fitstats, control-safety, and result-contract tests in 4.71 seconds plus git diff --check. The owner PR remains Draft, REVIEW_REQUIRED, and blocked; its hosted rollup is not green and independent approval is incomplete. Released item-covariate and CAT exposure-control APIs remain limited prerequisites, not executed buyer evidence.

Review remediation

  • automatic streaming now rejects a worker excluded by provider policy while preserving explicit requested-model behavior
  • persistence reuses one observation snapshot; propensity weighting reads the finalized selected candidate probability
  • evidence wording distinguishes winner probability 0.85 from the 0.05 exploration floor; duplicate DIF citation removed

Summary by CodeRabbit

  • 새로운 기능

    • 유사 컨텍스트의 정보를 활용하는 선택적 웜 스타트를 지원합니다.
    • 후보별 선택 근거와 라우팅 성능 증거를 더욱 상세히 기록합니다.
    • 심리측정 예측 품질과 후보 온보딩 보정 검증을 확장했습니다.
    • 스트리밍 라우팅에서 사용 가능한 작업자가 없을 때 명확한 오류를 표시합니다.
  • 문서

    • 심리측정 라우팅 검증 기준과 학술 참고문헌을 보강했습니다.
  • 버그 수정

    • 배포·정책 변경 시 이전 관측 데이터 재사용을 방지합니다.
    • 후보군 변경 및 동시성 상황에서 관측 데이터 정리를 안정화했습니다.
    • 미관측 질의와 후보 배포의 예측 결과를 구분해 보고합니다.

Hosted Ubuntu exposed platform-level floating drift in the synthetic item-covariate absolute-error literal. Commit ba010d5 asserts the published absolute-error calculation against its own estimated and true parameters while retaining the separate recovery-value assertion; the focused hosted-failure reproduction passes locally.

Signed-off-by: Seongho Bae <me@seonghobae.me>
Signed-off-by: Seongho Bae <me@seonghobae.me>
Signed-off-by: Seongho Bae <me@seonghobae.me>
Signed-off-by: Seongho Bae <me@seonghobae.me>
Signed-off-by: Seongho Bae <me@seonghobae.me>
Signed-off-by: Seongho Bae <me@seonghobae.me>
Signed-off-by: Seongho Bae <me@seonghobae.me>
Signed-off-by: Seongho Bae <me@seonghobae.me>
Signed-off-by: Seongho Bae <me@seonghobae.me>
Signed-off-by: Seongho Bae <me@seonghobae.me>
Signed-off-by: Seongho Bae <me@seonghobae.me>
Signed-off-by: Seongho Bae <me@seonghobae.me>
Record versioned deployment units and the invariance, DIF, uncertainty, judge, and adaptive-exposure gates required before fitted routing values can be treated as measurements.

Signed-off-by: Seongho Bae <me@seonghobae.me>
Normalize scholarly DOI and arXiv identifiers and require every citation used by tracked Python or Markdown to appear in the central paper register.

Signed-off-by: Seongho Bae <me@seonghobae.me>
Track ICLR, ACL Anthology, OpenReview, Anthropic research, and publisher DOI forms so the repository-wide paper register cannot silently omit non-arXiv sources.

Signed-off-by: Seongho Bae <me@seonghobae.me>
Keep two-neighbor interpolation experimental, bind observations to complete declared candidate configuration, and compute overflow-safe cosine similarity.

Signed-off-by: Seongho Bae <me@seonghobae.me>
Describe the experimental-only interpolation, deployment-bound evidence identity, and overflow-safe similarity behavior.

Signed-off-by: Seongho Bae <me@seonghobae.me>
Keep judge observations only while their complete deployment configuration remains active, both across restarts and runtime pool changes.

Signed-off-by: Seongho Bae <me@seonghobae.me>
Delete skipped deployment observations during reload so reverting an old configuration cannot revive invalid psychometric evidence.

Signed-off-by: Seongho Bae <me@seonghobae.me>
Preserve the validated single-neighbor production behavior and invalidate psychometric observations when the active role effort or sampling policy changes.

Signed-off-by: Seongho Bae <me@seonghobae.me>
Document the preserved production neighbor rule, decode-policy evidence identity, local verification, and remaining buyer-held-out validity gates.

Signed-off-by: Seongho Bae <me@seonghobae.me>
Compare single- and two-neighbor routing on paired held-out contexts and report deterministic bootstrap intervals for Brier, log-loss, and regret deltas.

Signed-off-by: Seongho Bae <me@seonghobae.me>
Record the exact source commit, confidence intervals, separate latency result, and closed production gate in operator guidance, ADR, and gap baseline.

Signed-off-by: Seongho Bae <me@seonghobae.me>
Measure 200 decisions per held-out context and bootstrap paired context-median latency differences instead of relying on one noisy timing sample.

Signed-off-by: Seongho Bae <me@seonghobae.me>
Search the bounded threshold grid and select the lowest p95 delay among candidates meeting the preregistered false-alarm and delay targets.

Signed-off-by: Seongho Bae <me@seonghobae.me>

Commit-Message-Assisted-by: Claude (via Claude Code)
Bind the bounded threshold search, target constraints, and improved p50 and p95 detection delay to its source commit.

Signed-off-by: Seongho Bae <me@seonghobae.me>

Commit-Message-Assisted-by: Claude (via Claude Code)
Select the threshold on calibration replications using a Wilson false-alarm bound, then report performance on an independent seed.

Signed-off-by: Seongho Bae <me@seonghobae.me>

Commit-Message-Assisted-by: Claude (via Claude Code)
Replace same-sample threshold claims with calibration-selection and independent held-out false-alarm and delay results.

Signed-off-by: Seongho Bae <me@seonghobae.me>

Commit-Message-Assisted-by: Claude (via Claude Code)
Record the missing discrimination orientation and latent-scale identification boundaries. Keep production admission closed until buyer evidence covers fit and uncertainty.

Signed-off-by: Seongho Bae <me@seonghobae.me>

Commit-Message-Assisted-by: Claude (via Claude Code)
Bind the local-dependence prerequisite to the current owner head and its detached focused audit without treating it as buyer or protected-main proof.

Signed-off-by: Seongho Bae <me@seonghobae.me>

Commit-Message-Assisted-by: Claude (via Claude Code)
Compare released maximum-information EAP calibration with a random query order on known synthetic candidates. Report query burden and prediction error without opening the buyer gate.

Signed-off-by: Seongho Bae <me@seonghobae.me>

Commit-Message-Assisted-by: Claude (via Claude Code)
Bind the maximum-information EAP screen to its research basis and measured query/error KPIs. Preserve the buyer and production validity boundaries.

Signed-off-by: Seongho Bae <me@seonghobae.me>

Commit-Message-Assisted-by: Claude (via Claude Code)
Pair candidate-level CAT outcomes with bootstrap intervals. Avoid claiming a theta-accuracy gain when its squared-error interval includes zero.

Signed-off-by: Seongho Bae <me@seonghobae.me>

Commit-Message-Assisted-by: Claude (via Claude Code)
Compare a bounded confidence-interval stopping rule with fixed-length candidate calibration. Report paired query and decision-accuracy intervals without changing production.

Signed-off-by: Seongho Bae <me@seonghobae.me>

Commit-Message-Assisted-by: Claude (via Claude Code)
Ground the bounded CI stop in classification-testing research and record its paired query and accuracy results without opening the buyer gate.

Signed-off-by: Seongho Bae <me@seonghobae.me>

Commit-Message-Assisted-by: Claude (via Claude Code)
Expose calibration burden and classification accuracy by distance from the decision cut so near-cut risk is not hidden by aggregate efficiency.

Signed-off-by: Seongho Bae <me@seonghobae.me>

Commit-Message-Assisted-by: Claude (via Claude Code)
Record distance-stratified query burden and accuracy so aggregate stopping efficiency cannot hide difficult boundary decisions.

Signed-off-by: Seongho Bae <me@seonghobae.me>

Commit-Message-Assisted-by: Claude (via Claude Code)
Signed-off-by: Seongho Bae <me@seonghobae.me>

Commit-Message-Assisted-by: Claude (via Claude Code)
Signed-off-by: Seongho Bae <me@seonghobae.me>

Commit-Message-Assisted-by: Claude (via Claude Code)
Signed-off-by: Seongho Bae <me@seonghobae.me>

Commit-Message-Assisted-by: Claude (via Claude Code)
Signed-off-by: Seongho Bae <me@seonghobae.me>

Commit-Message-Assisted-by: Claude (via Claude Code)
Signed-off-by: Seongho Bae <me@seonghobae.me>

Commit-Message-Assisted-by: Claude (via Claude Code)
Signed-off-by: Seongho Bae <me@seonghobae.me>

Commit-Message-Assisted-by: Claude (via Claude Code)
Signed-off-by: Seongho Bae <me@seonghobae.me>

Commit-Message-Assisted-by: Claude (via Claude Code)
Signed-off-by: Seongho Bae <me@seonghobae.me>

Commit-Message-Assisted-by: Claude (via Claude Code)
Signed-off-by: Seongho Bae <me@seonghobae.me>

Commit-Message-Assisted-by: Claude (via Claude Code)
Signed-off-by: Seongho Bae <me@seonghobae.me>

Commit-Message-Assisted-by: Claude (via Claude Code)
Signed-off-by: Seongho Bae <me@seonghobae.me>

Commit-Message-Assisted-by: Claude (via Claude Code)
Signed-off-by: Seongho Bae <me@seonghobae.me>

Commit-Message-Assisted-by: Claude (via Claude Code)
Signed-off-by: Seongho Bae <me@seonghobae.me>

Commit-Message-Assisted-by: Claude (via Claude Code)
seonghobae added a commit that referenced this pull request Sep 5, 2026
PR #1067 has exactly the same tree as #1058, already an ancestor of #1064.
Preserve both histories and the complete #1064 tree without reapplying
identical cherry-picked changes. The successor update remains a fast-forward.

Source tree: 8735f95
Predecessors: #1058, #1059, #1061, #1062, #1064.
No predecessor is closed before protected delivery and delta verification.

Signed-off-by: Seongho Bae <me@seonghobae.me>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant