feat: IRT classification accuracy and consistency (Rudner 2001/2005; Lee 2010 via cacIRT) - #223
Conversation
There was a problem hiding this comment.
Pull request overview
Adds IRT cut-score classification accuracy and consistency computations (Rudner normal-approximation; Lee summed-score via Lord–Wingersky) to the Rust core with PyO3 bindings and a validating Python wrapper, plus Rust/Python fixtures that pin results to an independent NumPy transcription (and targeted edge-case anchors like left-closed cuts).
Changes:
- Implement
mlsirm_core::classification::{rudner_classification, lee_classification}and expose results as a sharedClassificationResult. - Add PyO3 bindings in
fast-mlsirm-pyplus a thin Python validation/marshaling wrapper exported fromfast_mlsirm. - Add Rust unit tests and Python feature tests anchoring parity vs independent reference literals and trust-boundary rejections; document the feature in
CHANGELOG.md.
Reviewed changes
Copilot reviewed 8 out of 8 changed files in this pull request and generated 2 comments.
Show a summary per file
| File | Description |
|---|---|
| tests/unit/classification_tests.rs | New Rust fixtures covering Rudner/Lee parity vs independent literals, edge cut behavior, and degenerate-input rejections. |
| tests/test_paper_features.py | Adds Python-side parity + trust-boundary tests for the new classification wrappers. |
| python/fast_mlsirm/classification.py | New Python wrapper (rudner_classification, lee_classification) with validation/marshaling into a typed result dataclass. |
| python/fast_mlsirm/init.py | Re-exports the new classification APIs at the package top level. |
| crates/mlsirm-core/src/lib.rs | Registers the new classification module. |
| crates/mlsirm-core/src/classification.rs | New Rust implementation + internal test module wiring. |
| crates/fast-mlsirm-py/src/lib.rs | Adds PyO3 bindings returning dicts for both classification methods. |
| CHANGELOG.md | Documents the new classification feature, scope, and tested divergences from cacIRT. |
Comments suppressed due to low confidence (1)
crates/mlsirm-core/src/classification.rs:323
assembleallocatesm * n_pointsentries for conditional outputs, but this call site doesn't guardcutscores.len() * n_pointsagainstusizeoverflow. On extreme inputs this could panic during allocation. Add a checked multiplication before callingassembleso the function returns a structured error instead of panicking.
Ok(assemble(n_points, cutscores, &wn, masses, |i| ts[i]))
💡 Add Copilot custom instructions for smarter, more guided reviews. Learn how to get started.
|
@copilot Fix the code for all comments in this review thread. When a review comment includes a suggested change, apply the suggestion exactly. Do not make changes beyond what is described in the linked review thread. |
Done. Applied all three fixes in
Also fixed the error message in |
097dacd to
eba0445
Compare
7ef3bd5 to
4e47e5a
Compare
eba0445 to
fb2cd6e
Compare
4e47e5a to
fb50514
Compare
fb2cd6e to
95b856f
Compare
fb50514 to
eee3e09
Compare
eee3e09 to
ab8dc8a
Compare
…Lee 2010 via cacIRT) Rebased onto main after #222 squash merge; keeps only classification feature files. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
ab8dc8a to
090a2df
Compare
Rebased onto main after #223 squash merge; keeps only parallel-analysis feature files. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
…224) * feat: Horn parallel analysis for component retention (paran oracle) Rebased onto main after #223 squash merge; keeps only parallel-analysis feature files. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com> * fix(ksirt): guard ksirt_occ against zero-item panic Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com> --------- Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Iteration 6 of the autonomous paper-implementation loop. Stacked on #222 (DETECT); review only the last commit.
What
mlsirm_core::classification— IRT classification accuracy and consistency for cut-score decisions, with PyO3 bindings and a thin validating Python wrapper (fast_mlsirm.rudner_classification,fast_mlsirm.lee_classification).Rud.P(rowMeans) and quadrature weightsRud.D(weighted.mean).mlsirm_core::scoring::lord_wingerskyrecursion; raw cut c splits scores at ceil(c); a point's true category is the raw-score interval containing its expected true score. Consistency = sum of squared category masses (formula appears in neither Rudner paper; source: cacIRT).Documented divergences from the cacIRT oracle
Lee.Dalone is right-closed in R; differs only when a value lands exactly on a cut).Out of scope: polytomous Lee,
np.cac, MLE/SEM ability helpers, TOtable/kappa.Evidence
math.erf, own LW recursion, never imports this crate)erfc, err < 1.2e-7), Lee 1e-12#[ignore])cargo test -p mlsirm-core-k classification)Disclosed unkillable identities: m=1 simultaneous == per-cut (anchored by m=2 fixtures); uniform weights == plain mean (anchored by non-uniform unnormalized weights).
References (APA 7th)