feat(routing): make psychometric evidence fail closed - #1064
Open
seonghobae wants to merge 149 commits into
Open
feat(routing): make psychometric evidence fail closed#1064seonghobae wants to merge 149 commits into
seonghobae wants to merge 149 commits into
Conversation
Signed-off-by: Seongho Bae <me@seonghobae.me>
Signed-off-by: Seongho Bae <me@seonghobae.me>
Signed-off-by: Seongho Bae <me@seonghobae.me>
Signed-off-by: Seongho Bae <me@seonghobae.me>
Signed-off-by: Seongho Bae <me@seonghobae.me>
Signed-off-by: Seongho Bae <me@seonghobae.me>
Signed-off-by: Seongho Bae <me@seonghobae.me>
Signed-off-by: Seongho Bae <me@seonghobae.me>
Signed-off-by: Seongho Bae <me@seonghobae.me>
Signed-off-by: Seongho Bae <me@seonghobae.me>
…toresearch-psychometric-observe-20260904
Signed-off-by: Seongho Bae <me@seonghobae.me>
Signed-off-by: Seongho Bae <me@seonghobae.me>
Record versioned deployment units and the invariance, DIF, uncertainty, judge, and adaptive-exposure gates required before fitted routing values can be treated as measurements. Signed-off-by: Seongho Bae <me@seonghobae.me>
Normalize scholarly DOI and arXiv identifiers and require every citation used by tracked Python or Markdown to appear in the central paper register. Signed-off-by: Seongho Bae <me@seonghobae.me>
Track ICLR, ACL Anthology, OpenReview, Anthropic research, and publisher DOI forms so the repository-wide paper register cannot silently omit non-arXiv sources. Signed-off-by: Seongho Bae <me@seonghobae.me>
Keep two-neighbor interpolation experimental, bind observations to complete declared candidate configuration, and compute overflow-safe cosine similarity. Signed-off-by: Seongho Bae <me@seonghobae.me>
Describe the experimental-only interpolation, deployment-bound evidence identity, and overflow-safe similarity behavior. Signed-off-by: Seongho Bae <me@seonghobae.me>
Keep judge observations only while their complete deployment configuration remains active, both across restarts and runtime pool changes. Signed-off-by: Seongho Bae <me@seonghobae.me>
Delete skipped deployment observations during reload so reverting an old configuration cannot revive invalid psychometric evidence. Signed-off-by: Seongho Bae <me@seonghobae.me>
Preserve the validated single-neighbor production behavior and invalidate psychometric observations when the active role effort or sampling policy changes. Signed-off-by: Seongho Bae <me@seonghobae.me>
Document the preserved production neighbor rule, decode-policy evidence identity, local verification, and remaining buyer-held-out validity gates. Signed-off-by: Seongho Bae <me@seonghobae.me>
Compare single- and two-neighbor routing on paired held-out contexts and report deterministic bootstrap intervals for Brier, log-loss, and regret deltas. Signed-off-by: Seongho Bae <me@seonghobae.me>
Record the exact source commit, confidence intervals, separate latency result, and closed production gate in operator guidance, ADR, and gap baseline. Signed-off-by: Seongho Bae <me@seonghobae.me>
Measure 200 decisions per held-out context and bootstrap paired context-median latency differences instead of relying on one noisy timing sample. Signed-off-by: Seongho Bae <me@seonghobae.me>
Search the bounded threshold grid and select the lowest p95 delay among candidates meeting the preregistered false-alarm and delay targets. Signed-off-by: Seongho Bae <me@seonghobae.me> Commit-Message-Assisted-by: Claude (via Claude Code)
Bind the bounded threshold search, target constraints, and improved p50 and p95 detection delay to its source commit. Signed-off-by: Seongho Bae <me@seonghobae.me> Commit-Message-Assisted-by: Claude (via Claude Code)
Select the threshold on calibration replications using a Wilson false-alarm bound, then report performance on an independent seed. Signed-off-by: Seongho Bae <me@seonghobae.me> Commit-Message-Assisted-by: Claude (via Claude Code)
Replace same-sample threshold claims with calibration-selection and independent held-out false-alarm and delay results. Signed-off-by: Seongho Bae <me@seonghobae.me> Commit-Message-Assisted-by: Claude (via Claude Code)
Record the missing discrimination orientation and latent-scale identification boundaries. Keep production admission closed until buyer evidence covers fit and uncertainty. Signed-off-by: Seongho Bae <me@seonghobae.me> Commit-Message-Assisted-by: Claude (via Claude Code)
Bind the local-dependence prerequisite to the current owner head and its detached focused audit without treating it as buyer or protected-main proof. Signed-off-by: Seongho Bae <me@seonghobae.me> Commit-Message-Assisted-by: Claude (via Claude Code)
Compare released maximum-information EAP calibration with a random query order on known synthetic candidates. Report query burden and prediction error without opening the buyer gate. Signed-off-by: Seongho Bae <me@seonghobae.me> Commit-Message-Assisted-by: Claude (via Claude Code)
Bind the maximum-information EAP screen to its research basis and measured query/error KPIs. Preserve the buyer and production validity boundaries. Signed-off-by: Seongho Bae <me@seonghobae.me> Commit-Message-Assisted-by: Claude (via Claude Code)
Pair candidate-level CAT outcomes with bootstrap intervals. Avoid claiming a theta-accuracy gain when its squared-error interval includes zero. Signed-off-by: Seongho Bae <me@seonghobae.me> Commit-Message-Assisted-by: Claude (via Claude Code)
Compare a bounded confidence-interval stopping rule with fixed-length candidate calibration. Report paired query and decision-accuracy intervals without changing production. Signed-off-by: Seongho Bae <me@seonghobae.me> Commit-Message-Assisted-by: Claude (via Claude Code)
Ground the bounded CI stop in classification-testing research and record its paired query and accuracy results without opening the buyer gate. Signed-off-by: Seongho Bae <me@seonghobae.me> Commit-Message-Assisted-by: Claude (via Claude Code)
Expose calibration burden and classification accuracy by distance from the decision cut so near-cut risk is not hidden by aggregate efficiency. Signed-off-by: Seongho Bae <me@seonghobae.me> Commit-Message-Assisted-by: Claude (via Claude Code)
Record distance-stratified query burden and accuracy so aggregate stopping efficiency cannot hide difficult boundary decisions. Signed-off-by: Seongho Bae <me@seonghobae.me> Commit-Message-Assisted-by: Claude (via Claude Code)
Signed-off-by: Seongho Bae <me@seonghobae.me> Commit-Message-Assisted-by: Claude (via Claude Code)
Signed-off-by: Seongho Bae <me@seonghobae.me> Commit-Message-Assisted-by: Claude (via Claude Code)
Signed-off-by: Seongho Bae <me@seonghobae.me> Commit-Message-Assisted-by: Claude (via Claude Code)
Signed-off-by: Seongho Bae <me@seonghobae.me> Commit-Message-Assisted-by: Claude (via Claude Code)
Signed-off-by: Seongho Bae <me@seonghobae.me> Commit-Message-Assisted-by: Claude (via Claude Code)
Signed-off-by: Seongho Bae <me@seonghobae.me> Commit-Message-Assisted-by: Claude (via Claude Code)
Signed-off-by: Seongho Bae <me@seonghobae.me> Commit-Message-Assisted-by: Claude (via Claude Code)
Signed-off-by: Seongho Bae <me@seonghobae.me> Commit-Message-Assisted-by: Claude (via Claude Code)
Signed-off-by: Seongho Bae <me@seonghobae.me> Commit-Message-Assisted-by: Claude (via Claude Code)
Signed-off-by: Seongho Bae <me@seonghobae.me> Commit-Message-Assisted-by: Claude (via Claude Code)
Signed-off-by: Seongho Bae <me@seonghobae.me> Commit-Message-Assisted-by: Claude (via Claude Code)
Signed-off-by: Seongho Bae <me@seonghobae.me> Commit-Message-Assisted-by: Claude (via Claude Code)
Signed-off-by: Seongho Bae <me@seonghobae.me> Commit-Message-Assisted-by: Claude (via Claude Code)
seonghobae
added a commit
that referenced
this pull request
Sep 5, 2026
PR #1067 has exactly the same tree as #1058, already an ancestor of #1064. Preserve both histories and the complete #1064 tree without reapplying identical cherry-picked changes. The successor update remains a fast-forward. Source tree: 8735f95 Predecessors: #1058, #1059, #1061, #1062, #1064. No predecessor is closed before protected delivery and delta verification. Signed-off-by: Seongho Bae <me@seonghobae.me>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Successor tracking
The complete delta at this PR's head
ba010d5a023f69204df73e0a8e4d9b0a5bf7ee2fis preserved in trusted-branch successor #1067 at1481c595dc1d16e7bf4b65addaf0bd30322cf2b8. Commit ancestry and tree equality are recorded indocs/product-technical-gap-baseline.md. Review and required-workflow integration continue there. This PR remains open until protected delivery and a fresh full-delta audit; later changes here must also be carried forward.Summary
Exact-head evidence
ba010d5a023f69204df73e0a8e4d9b0a5bf7ee2f200 passedon predecessor head32 passed in 15.50s3417 passed, 2 skipped in 666.99sgit diff --check: passMeasurement boundary
The local synthetic benchmark preserves its accuracy improvement, but its latency confidence interval does not authorize production. Live deterministic routing records candidate, attempt, selection, and policy identity with
propensity_status=not_identified; no counterfactual correction is claimed.A fixed-seed benchmark-only 20% epsilon-greedy design assigns all four candidates probability at least 0.05 over 24,000 trials. Naively averaging the adaptively selected observations gives RMSE 0.321979 against known truth; logged inverse-propensity estimation reduces RMSE to 0.008943, an improvement of 0.313036, and all four 95% intervals cover known truth. Six-anchor Stocking-Lord linking recovers slope 1.3 and intercept -0.4 with true-parameter RMSE 3.24e-16. Purified logistic DIF detects the one injected cohort shift with recall 1.0 and zero false positives. A connected many-facet Rasch fit recovers three judge severities with RMSE 0.018292.
Two-neighbor interpolation reduces held-out logit calibration RMSE from
0.231235to0.025803; the paired candidate-minus-baseline 95% interval is[-0.208030, -0.202924]. Calibration slope moves from0.991445to1.014030. These known synthetic probabilities do not establish calibration on buyer outcomes.The released item-side covariate fit converges after 941 iterations and estimates
-0.789650for truedelta=-0.8(absolute error0.010350). The released Oakes-information API gives item-intercept RMSE0.039160, 95% interval coverage1.0, and mean interval width0.295945for six known intercepts. It conditions on population parameters and rejects anchors, zero inflation, and item covariates. A separate recalibration-invariance screen links through seven stable anchors and flags the one injected drift item with no stable-item false positive; its fixed0.25tolerance is an effect-size rule, not a significance test. Nonparametric person fit ranks one injected inverted candidate pattern first among 1,000 patterns with ZU3 separation1.818719, but does not infer a cause or universal cutoff. Horn parallel analysis retains both known dimensions in 1,000 binary response vectors across 12 items; because Pearson-PCA screening is not construct identification or confirmatory fit, buyer dimensionality remains unexecuted. Limited-information M2 accepts the fitted one-factor synthetic case atp=0.105619and detects known two-factor misspecification atp≈2.27e-41; sensitivity varies by misspecification, so global buyer model fit stays unexecuted. Posterior empirical reliability rises from0.366437to0.800436when true item discrimination rises from0.45to1.5; reliability remains separate from model fit and validity. Spreading 12 item difficulties across[-2, 2]improves worst information over trait points[-2, 0, 2]by15.789%and lowers worst conditional standard error from0.890897to0.827931, while exposing the center-precision tradeoff. At a declared synthetic cut of zero, lowering score standard error from0.8to0.2raises expected classification accuracy from0.814182to0.996895and consistency from0.710275to0.993829; no buyer cut or asymmetric decision cost is inferred. These synthetic checks validate calculation contracts only; they do not establish buyer outcomes, live randomization, invariant scales, buyer uncertainty, construct validity, or production authorization.A candidate-roster screen independently calibrates 20- and 16-candidate synthetic rosters and links 200 common items. The retained candidates have linked-score RMSE
0.010866, correlation0.999999, and maximum shift0.016498. Versioned buyer rosters and preregistered shift targets remain absent, so this does not establish sample-free buyer measurement.A two-facet G-study separates candidate, query, occasion, and interaction variance in an 80-candidate, 12-query, four-occasion synthetic tensor. The D-study raises dependability from
0.401565for one query and one occasion to0.730184for six and two and0.849616for 12 and four. A complete balanced buyer design, random-facet justification, and a registered target remain absent.A decision-utility screen raises synthetic predictive validity from
0.2to0.6, increasing Taylor-Russell selected success from0.500273to0.723515and Brogden-Cronbach-Gleser net utility from2,042.21to5,626.64. Holding validity fixed while increasing total measurement cost from2,000to10,000leaves selected success unchanged but makes net utility negative at-2,373.36. This personnel-selection model is an analogue, not validated routing economics; buyer-valued outcomes, actual costs and volume, selection ratio, and a preregistered utility target remain absent.The predictive-fit report now distinguishes held-out queries for known candidate deployments from a held-out candidate deployment. Existing synthetic Brier and log-loss results belong only to the first task. Executing the unseen-candidate path across 24 contexts produces zero psychometric prediction coverage because the router does not fabricate an unobserved candidate score. This measured failure prevents query holdout evidence from being presented as cold-start deployment generalization.
A separate few-observation onboarding screen reuses released Rust-backed maximum-information EAP selection across 400 known synthetic candidates and a 31-query bank. It reaches target SE 0.5 after 7.1775 queries on average versus 10.47 for a seeded random order. Paired 95% intervals are
[-3.4125, -3.18]queries,[-0.080874, 0.006129]theta squared error, and[-0.008804, -0.005716]unobserved-probability squared error. The theta interval includes zero, so no general theta-accuracy gain is claimed. The maximum 12 calibration calls measure onboarding burden, not live routing latency; zero-observation coverage and production authorization remain unchanged.A classification-oriented screen stops when a 95% normal interval excludes a declared zero cut, with the same 12-query maximum. Across 400 known synthetic candidates, it stops early for 41%, averages 9.875 queries, and exactly matches the fixed-length decisions and 0.9125 accuracy. Paired intervals are
[-2.425, -1.835]queries and[0, 0]accuracy. Buyer cuts and costs, near-cut risk, interval calibration, and live provider latency remain absent, so this does not authorize production.Distance-from-cut strata expose the aggregate limit. Within 0.5 of the synthetic cut, candidates stop early only 3%, average 11.86 queries, and reach 0.70 accuracy. Candidates at least 1.0 away stop early 68%, average 8.305 queries, and reach 1.0 accuracy. These descriptive known-truth strata are not calibrated buyer subgroup guarantees.
Keeping unresolved cases explicit changes the interpretation: 42.5% of candidates are confidence-resolved and all resolved synthetic decisions are correct, while only 3% of near-cut candidates resolve. The remaining cases stay unresolved rather than being treated as production-ready decisions; buyer fallback policy and calibrated coverage remain absent.
A reject-option frontier now separates threshold selection from evaluation. Maximizing development coverage subject to a 2.5% Wilson 95% selective-risk upper bound selects
z=1.645. Againstz=1.96on the same independent responses, coverage rises from 44.25% to 56% and all-candidate mean queries fall from 9.88 to 8.395. Paired 95% intervals are[8.75, 14.75]percentage points and[-1.715, -1.2625]queries. Observed selective risk is zero with a 1.686% Wilson upper bound, and directional coverage differs by 3 points. A ten-seed replication audit rejects this threshold despite the single-run efficiency. Coverage gain remains 9.75–15 percentage points and all-candidate query reduction remains 1.27–1.485, but selective risk reaches 1.802% and its Wilson upper bound reaches 4.540%. Only 20% of replications satisfy the declared 2.5% ceiling, soz=1.645is markedrejected_not_replication_stableand cannot become a production threshold. The ten-run pass-rate MCSE is 0.1265; coverage-delta, all-candidate query-delta, and selective-risk MCSEs are 0.00619, 0.02257, and 0.00211. A conservative binomial design requires 400 replications for target pass-rate MCSE 0.025, so ten runs are a falsification screen rather than a precise operating-characteristic estimate.A stdlib-only hot-path change replaces generator iteration in cosine dot products with map(operator.mul, ...), preserving math.fsum and multiplication order. Across ten fresh-process before/after runs, candidate p50 latency falls from a median 0.01325 ms to 0.007708 ms (41.83%), and the paired candidate-minus-baseline latency-CI upper bound falls from a median 0.000910 ms to 0.000635 ms (30.14%). The bound remains positive, so the decision-latency gate stays failed. These local CPU timings do not establish end-to-end provider latency.
A functional-impact screen also identifies the one item producing TCC-area difference
0.123355and reduces the residual difference to zero. Its backward-only fixed-threshold search is narrower than Guo, Zheng, and Chang's complete stepwise method and does not supply a buyer cutoff.A one-stream Bernoulli CUSUM screen measures the false-alarm and temporal-detection tradeoff separately. A bounded search over 11 thresholds uses 500 calibration runs and a 95% Wilson false-alarm upper bound, then evaluates selected threshold
6.6on an independent 500-run seed. Held-out false alarms are2.4%with upper bound4.15%; detection-delay p50 is 10 and p95 is 20. These are synthetic calculation contracts, not Chen, Lee, and Li's multistream Bayesian compound-risk procedure or buyer-approved thresholds.Alternate-form linear equating recovers a known slope
2and intercept1, reducing raw cross-form RMSE from6.782330to zero. Three hundred bootstrap 95% intervals cover all 11 known equivalent scores, but equal synthetic form populations do not establish buyer score comparability.Owner dependency
Local-independence indices remain proposed in ContextualWisdomLab/fast-mlsirm#1748 at exact head
8461914a5bf04f9732add77761dd121bbec00103. A detached exact-head audit passes 44 focused fitstats, control-safety, and result-contract tests in 4.71 seconds plusgit diff --check. The owner PR remains Draft,REVIEW_REQUIRED, and blocked; its hosted rollup is not green and independent approval is incomplete. Released item-covariate and CAT exposure-control APIs remain limited prerequisites, not executed buyer evidence.Review remediation
Summary by CodeRabbit
새로운 기능
문서
버그 수정
Hosted Ubuntu exposed platform-level floating drift in the synthetic item-covariate absolute-error literal. Commit ba010d5 asserts the published absolute-error calculation against its own estimated and true parameters while retaining the separate recovery-value assertion; the focused hosted-failure reproduction passes locally.