Skip to content

feat: IRT classification accuracy and consistency (Rudner 2001/2005; Lee 2010 via cacIRT) - #223

Merged
seonghobae merged 1 commit into
mainfrom
seonghobae-classification-accuracy
Jul 25, 2026
Merged

seonghobae merged 1 commit into
mainfrom
seonghobae-classification-accuracy

Conversation

@seonghobae

Copy link
Copy Markdown
Contributor

Iteration 6 of the autonomous paper-implementation loop. Stacked on #222 (DETECT); review only the last commit.

What

mlsirm_core::classification — IRT classification accuracy and consistency for cut-score decisions, with PyO3 bindings and a thin validating Python wrapper (fast_mlsirm.rudner_classification, fast_mlsirm.lee_classification).

  • Rudner normal-approximation method (Rudner, 2001, 2005 — both read in full via PARE archives): observed score at ability theta modeled as N(theta, sem^2); per-cut and simultaneous (m+1-category) accuracy; conditional and marginal outputs; weights normalized internally so uniform weights reproduce cacIRT Rud.P (rowMeans) and quadrature weights Rud.D (weighted.mean).
  • Lee summed-score method (Lee, 2010, as cited in Lathrop's CRAN cacIRT 1.4 sources, read line by line — the paper itself is paywalled): exact summed-score distribution per point via the existing mlsirm_core::scoring::lord_wingersky recursion; raw cut c splits scores at ceil(c); a point's true category is the raw-score interval containing its expected true score. Consistency = sum of squared category masses (formula appears in neither Rudner paper; source: cacIRT).

Documented divergences from the cacIRT oracle

  1. Left-closed category intervals everywhere (Lee.D alone is right-closed in R; differs only when a value lands exactly on a cut).
  2. Item probabilities must lie strictly in (0, 1): P == 1 breaks the R oracle's hazard recursion; rejecting P == 0 is a stricter policy than the oracle, labeled as such.
  3. Simultaneous outputs always populated (cacIRT emits them only for 2+ cuts); for one cut they equal the per-cut values — identity anchored by m=2 fixtures.
  4. Raw cuts restricted to (0, n_items] with strictly increasing ceil-mapped boundaries (a reversed-slice hazard in the R oracle).

Out of scope: polytomous Lee, np.cac, MLE/SEM ability helpers, TOtable/kappa.

Evidence

Check Result
Fixture parity vs independent NumPy transcription (exact math.erf, own LW recursion, never imports this crate) Rudner 1e-6 (crate erfc, err < 1.2e-7), Lee 1e-12
Left-closed anchors Rudner theta exactly on cut −0.4; Lee dyadic true score exactly 3.0 on raw cut 3.0
Mutation spot-checks (all killed, restored) M1 right-closed categorization: 3 tests FAIL. M2 consistency without squaring: 3 FAIL. M3 floor-vs-ceil raw cut (non-integer cut 2.4): 1 FAIL. M4 skipped weight normalization (weights sum 8.5): 2 FAIL
500-rep Monte Carlo (#[ignore]) informative 40-item tests dominate noisy 8-item tests in simultaneous accuracy/consistency in >95% of reps — passed
cargo test -p mlsirm-core 470 passed, 0 failed
Python suite (-k classification) 2 passed
Adversarial impl-review 1 BLOCKING defect found and fixed with regression tests: finite weights summing to inf normalized to all-zero and silently zeroed marginals; now rejected ("weights sum overflows f64")

Disclosed unkillable identities: m=1 simultaneous == per-cut (anchored by m=2 fixtures); uniform weights == plain mean (anchored by non-uniform unnormalized weights).

References (APA 7th)

  • Lathrop, Q. N. (2015). cacIRT: Classification accuracy and consistency under item response theory (Version 1.4) [R package]. CRAN.
  • Lee, W.-C. (2010). Classification consistency and accuracy for complex assessments using item response theory. Journal of Educational Measurement, 47(1), 1-17. (As cited in Lathrop, 2015; not read.)
  • Lord, F. M., & Wingersky, M. S. (1984). Comparison of IRT true-score and equipercentile observed-score "equatings". Applied Psychological Measurement, 8(4), 453-461.
  • Rudner, L. M. (2001). Computing the expected proportions of misclassified examinees. PARE, 7(14). https://doi.org/10.7275/an9m-2035
  • Rudner, L. M. (2005). Expected classification accuracy. PARE, 10(13). https://doi.org/10.7275/56a5-6b14

Copilot AI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

Adds IRT cut-score classification accuracy and consistency computations (Rudner normal-approximation; Lee summed-score via Lord–Wingersky) to the Rust core with PyO3 bindings and a validating Python wrapper, plus Rust/Python fixtures that pin results to an independent NumPy transcription (and targeted edge-case anchors like left-closed cuts).

Changes:

  • Implement mlsirm_core::classification::{rudner_classification, lee_classification} and expose results as a shared ClassificationResult.
  • Add PyO3 bindings in fast-mlsirm-py plus a thin Python validation/marshaling wrapper exported from fast_mlsirm.
  • Add Rust unit tests and Python feature tests anchoring parity vs independent reference literals and trust-boundary rejections; document the feature in CHANGELOG.md.

Reviewed changes

Copilot reviewed 8 out of 8 changed files in this pull request and generated 2 comments.

Show a summary per file
File Description
tests/unit/classification_tests.rs New Rust fixtures covering Rudner/Lee parity vs independent literals, edge cut behavior, and degenerate-input rejections.
tests/test_paper_features.py Adds Python-side parity + trust-boundary tests for the new classification wrappers.
python/fast_mlsirm/classification.py New Python wrapper (rudner_classification, lee_classification) with validation/marshaling into a typed result dataclass.
python/fast_mlsirm/init.py Re-exports the new classification APIs at the package top level.
crates/mlsirm-core/src/lib.rs Registers the new classification module.
crates/mlsirm-core/src/classification.rs New Rust implementation + internal test module wiring.
crates/fast-mlsirm-py/src/lib.rs Adds PyO3 bindings returning dicts for both classification methods.
CHANGELOG.md Documents the new classification feature, scope, and tested divergences from cacIRT.
Comments suppressed due to low confidence (1)

crates/mlsirm-core/src/classification.rs:323

  • assemble allocates m * n_points entries for conditional outputs, but this call site doesn't guard cutscores.len() * n_points against usize overflow. On extreme inputs this could panic during allocation. Add a checked multiplication before calling assemble so the function returns a structured error instead of panicking.
    Ok(assemble(n_points, cutscores, &wn, masses, |i| ts[i]))

💡 Add Copilot custom instructions for smarter, more guided reviews. Learn how to get started.

Comment thread crates/mlsirm-core/src/classification.rs
Comment thread crates/mlsirm-core/src/classification.rs Outdated
@seonghobae

Copy link
Copy Markdown
Contributor Author

@copilot Fix the code for all comments in this review thread.

When a review comment includes a suggested change, apply the suggestion exactly.

Do not make changes beyond what is described in the linked review thread.

Copilot AI commented Jul 24, 2026

Copy link
Copy Markdown
Contributor

@copilot Fix the code for all comments in this review thread.

When a review comment includes a suggested ch...

Done. Applied all three fixes in crates/mlsirm-core/src/classification.rs:

  1. rudner_classification (line ~234): added cutscores.len().checked_mul(n).ok_or_else(…)? guard before assemble to catch m * n overflow before the allocation.
  2. lee_classification overflow guards (lines ~300-313):
    • Replaced the unchecked n_items + 1 with n_items.checked_add(1).ok_or_else(…)? to get sc.
    • Replaced the unchecked n_points * sc allocation with a n_points.checked_mul(sc).ok_or_else(…)? dist_len, then vec![0.0; dist_len].
    • Added the same cutscores.len().checked_mul(n_points) guard before assemble.

Also fixed the error message in rudner_classification to say "n" (matching the actual variable) instead of "n_points", and added a comment explaining the guards are solely to detect overflow before assemble allocates m * n entries. All 6 classification tests pass.

Copilot AI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

Copilot reviewed 8 out of 8 changed files in this pull request and generated no new comments.

@seonghobae
seonghobae force-pushed the seonghobae-detect-dimensionality branch from 097dacd to eba0445 Compare July 25, 2026 09:27
@seonghobae
seonghobae force-pushed the seonghobae-classification-accuracy branch from 7ef3bd5 to 4e47e5a Compare July 25, 2026 09:29
@seonghobae
seonghobae force-pushed the seonghobae-detect-dimensionality branch from eba0445 to fb2cd6e Compare July 25, 2026 10:02
@seonghobae
seonghobae force-pushed the seonghobae-classification-accuracy branch from 4e47e5a to fb50514 Compare July 25, 2026 10:03
@seonghobae
seonghobae force-pushed the seonghobae-detect-dimensionality branch from fb2cd6e to 95b856f Compare July 25, 2026 10:13
@seonghobae
seonghobae force-pushed the seonghobae-classification-accuracy branch from fb50514 to eee3e09 Compare July 25, 2026 10:14
Base automatically changed from seonghobae-detect-dimensionality to main July 25, 2026 10:19
@seonghobae
seonghobae force-pushed the seonghobae-classification-accuracy branch from eee3e09 to ab8dc8a Compare July 25, 2026 10:21
Copilot AI review requested due to automatic review settings July 25, 2026 10:29
…Lee 2010 via cacIRT)

Rebased onto main after #222 squash merge; keeps only classification feature files.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>

Copilot AI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

Copilot reviewed 8 out of 8 changed files in this pull request and generated 1 comment.

Comment thread crates/mlsirm-core/src/classification.rs
@seonghobae
seonghobae merged commit 337e520 into main Jul 25, 2026
33 checks passed
@seonghobae
seonghobae deleted the seonghobae-classification-accuracy branch July 25, 2026 10:36
seonghobae added a commit that referenced this pull request Jul 25, 2026
Rebased onto main after #223 squash merge; keeps only parallel-analysis feature files.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
seonghobae added a commit that referenced this pull request Jul 25, 2026
…224)

* feat: Horn parallel analysis for component retention (paran oracle)

Rebased onto main after #223 squash merge; keeps only parallel-analysis feature files.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>

* fix(ksirt): guard ksirt_occ against zero-item panic

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>

---------

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants