Skip to content

feat(model-selection): govern structural pair decisions - #987

Closed
seonghobae wants to merge 11 commits into
mainfrom
codex/structural-selection-governor-608
Closed

seonghobae wants to merge 11 commits into
mainfrom
codex/structural-selection-governor-608

Conversation

@seonghobae

Copy link
Copy Markdown
Contributor

Advances #608 with one bounded governance slice on current protected main.

Scope

  • Keep factor retention separate from structural model selection.
  • Consume explicit ModelRelationEvidence; never infer nesting/overlap from model names.
  • Require the relation-appropriate LR/bootstrap/Vuong stage before admitting an already-computed pairwise outcome.
  • Gate any winner on separate recovery and intended-score interpretation evidence.
  • Prefer the simpler candidate only when an already-governed comparison establishes practical equivalence and its score interpretation is supported.
  • Return explicit no-winner states for unknown relation, indistinguishability, insufficient recovery, and unsupported score interpretation.

Python performs validation and policy orchestration only. This PR adds no likelihood, bootstrap, Vuong, predictive, recovery, scoreability, or practical-equivalence arithmetic; production numerical psychometrics remains Rust-owned.

Test-first lineage

  • 9a0083ffe4cc95dccb30dbe45882351f9becd32e — RED public contract/tests before the module existed.
  • b7da356b94eb469dd0547b4ef16697d9289185c5 — GREEN governor implementation.
  • 7ade44dde9c2b6a7ac0732f9f6e80f1e7546de9a — changelog trace.
  • b10726ff89f4f8432379273298eb51a42c518174 — scientific interpretation/APA 7 doctoring.

Keep Draft until exact-current-head hosted CI/security/review evidence is terminal. Predecessor evidence does not transfer.

@coderabbitai

coderabbitai Bot commented Aug 18, 2026 •

Copy link
Copy Markdown
Contributor

Review Change Stack

Warning

Review limit reached

@seonghobae, you've reached your PR review limit, so we couldn't start this review.

Next review available in: 7 minutes

Limit details: You’ve used all 1 included review currently available under your plan.

You've used all free OSS reviews for now. Wait for the free limit to reset to keep reviewing this public repository.

How can I continue?

After more reviews become available, a review can be triggered using the @coderabbitai review command as a PR comment. Alternatively, push new commits to this PR.

To avoid repeated limits, reduce automatic review volume by pausing incremental auto-reviews earlier, using label-based review opt-in, excluding WIP or generated PR titles, or requesting reviews manually when the PR is ready. If your team needs uninterrupted high-volume reviews, an organization admin can enable usage-based reviews.

How do review limits work?

CodeRabbit enforces per-developer PR review limits for each organization. Most developers receive the normal plan review availability.

For paid Pro and Pro+ PR reviews, CodeRabbit uses adaptive limits for sustained high-volume activity. When a developer's recent PR review activity reaches the 95th percentile or higher among CodeRabbit users, additional reviews become available more gradually as earlier reviews age out of the rolling window.

Please refer docs for additional details.

Review details
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: CHILL

Plan: Pro Plus

Run ID: d8f39201-b04b-4269-86e2-a8917d17ce41

📥 Commits

Reviewing files that changed from the base of the PR and between 04d0bc2 and 42b849c.

📒 Files selected for processing (7)
  • docs/changelog.d/608-structural-selection-governor.md
  • docs/doctoring/structural_selection_governor.md
  • python/fast_mlsirm/model_relation.py
  • python/fast_mlsirm/structural_selection.py
  • tests/test_model_relation_contract.py
  • tests/test_structural_selection_governor.py
  • tests/test_structural_selection_governor_edges.py
✨ Finishing Touches
📝 Generate docstrings
  • Create stacked PR
  • Commit on current branch
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Commit unit tests in branch codex/structural-selection-governor-608

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@seonghobae
seonghobae marked this pull request as ready for review August 18, 2026 03:39
@seonghobae
seonghobae enabled auto-merge (squash) August 18, 2026 03:40
@seonghobae
seonghobae marked this pull request as draft August 18, 2026 04:10
auto-merge was automatically disabled August 18, 2026 04:10

Pull request was converted to draft

@seonghobae
seonghobae marked this pull request as ready for review August 18, 2026 05:00
@seonghobae
seonghobae enabled auto-merge (squash) August 18, 2026 05:01

@opencode-agent opencode-agent Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

OpenCode cannot approve yet because required coverage evidence did not pass.

Review outcome

1. HIGH .github/workflows/opencode-review.yml:1 - Coverage evidence did not prove required test/docstring evidence

  • Problem: The required coverage-evidence job result was failure, so OpenCode cannot establish approval sufficiency for this head.

  • Root cause: Automated approval is only valid when the same-head coverage-evidence job proves supported repository test suites passed and configured docstring gates passed or were advisory, or reports not applicable because no supported source files or package manifests exist. Missing, failed, skipped, unavailable, or unsupported-tooling test evidence is a blocker.

  • Fix: Install or configure the repository test/docstring evidence tooling when source files or package manifests exist, rerun the current-head coverage-evidence job, and approve only after it reports success with required evidence or explicit no-source not-applicable evidence.

  • Regression test: Keep the approval branch checking needs.coverage-evidence.result == success before posting APPROVE, and publish REQUEST_CHANGES when coverage-evidence blocker states such as cancelled, skipped, failed, unsupported-tooling, or below-100 evidence are present.

  • Result: REQUEST_CHANGES

  • Reason: coverage-evidence result was failure, so required test/docstring evidence was not proven for current head 42b849c4c05c4464205c3afe71e39ce8c0cdae2a.

  • Head SHA: 42b849c4c05c4464205c3afe71e39ce8c0cdae2a

  • Workflow run: 32124705657

  • Workflow attempt: 1

Coverage evidence

Coverage evidence job did not run or did not publish coverage evidence.

Changed-File Evidence Map

flowchart LR
  PR["PR changed files"] --> Evidence["OpenCode bounded evidence"]
  Evidence --> S1["Docs (2 files)"]
  S1 --> I1["operator or user guidance"]
  I1 --> R1["Review risk: Docs (2 files)"]
  R1 --> V1["docs review"]
  Evidence --> S2["Changed file (2 files)"]
  S2 --> I2["repository behavior"]
  I2 --> R2["Review risk: Changed file (2 files)"]
  R2 --> V2["required checks"]
  Evidence --> S3["Test (3 files)"]
  S3 --> I3["regression suite"]
  I3 --> R3["Review risk: Test (3 files)"]
  R3 --> V3["targeted test run"]
Loading

@opencode-agent

Copy link
Copy Markdown
Contributor

OpenCode Review Overview

  • Head SHA: 42b849c4c05c4464205c3afe71e39ce8c0cdae2a
  • Workflow run: 32124705657
  • Workflow attempt: 1
  • Gate result: REQUEST_CHANGES (approval step)

Pull request overview

OpenCode cannot approve yet because required coverage evidence did not pass.

Review outcome

1. HIGH .github/workflows/opencode-review.yml:1 - Coverage evidence did not prove required test/docstring evidence

  • Problem: The required coverage-evidence job result was failure, so OpenCode cannot establish approval sufficiency for this head.

  • Root cause: Automated approval is only valid when the same-head coverage-evidence job proves supported repository test suites passed and configured docstring gates passed or were advisory, or reports not applicable because no supported source files or package manifests exist. Missing, failed, skipped, unavailable, or unsupported-tooling test evidence is a blocker.

  • Fix: Install or configure the repository test/docstring evidence tooling when source files or package manifests exist, rerun the current-head coverage-evidence job, and approve only after it reports success with required evidence or explicit no-source not-applicable evidence.

  • Regression test: Keep the approval branch checking needs.coverage-evidence.result == success before posting APPROVE, and publish REQUEST_CHANGES when coverage-evidence blocker states such as cancelled, skipped, failed, unsupported-tooling, or below-100 evidence are present.

  • Result: REQUEST_CHANGES

  • Reason: coverage-evidence result was failure, so required test/docstring evidence was not proven for current head 42b849c4c05c4464205c3afe71e39ce8c0cdae2a.

  • Head SHA: 42b849c4c05c4464205c3afe71e39ce8c0cdae2a

  • Workflow run: 32124705657

  • Workflow attempt: 1

Coverage evidence

Coverage evidence job did not run or did not publish coverage evidence.

Changed-File Evidence Map

flowchart LR
  PR["PR changed files"] --> Evidence["OpenCode bounded evidence"]
  Evidence --> S1["Docs (2 files)"]
  S1 --> I1["operator or user guidance"]
  I1 --> R1["Review risk: Docs (2 files)"]
  R1 --> V1["docs review"]
  Evidence --> S2["Changed file (2 files)"]
  S2 --> I2["repository behavior"]
  I2 --> R2["Review risk: Changed file (2 files)"]
  R2 --> V2["required checks"]
  Evidence --> S3["Test (3 files)"]
  S3 --> I3["regression suite"]
  I3 --> R3["Review risk: Test (3 files)"]
  R3 --> V3["targeted test run"]
Loading

@opencode-agent
opencode-agent Bot disabled auto-merge August 18, 2026 10:58
@seonghobae
seonghobae enabled auto-merge (squash) August 18, 2026 12:03
@opencode-agent
opencode-agent Bot disabled auto-merge August 18, 2026 12:44

Copy link
Copy Markdown
Contributor Author

Current-head RCA for 42b849c4c05c4464205c3afe71e39ce8c0cdae2a: the formal OpenCode CHANGES_REQUESTED is still authoritative, but the first causal boundary in central run 32124705657 is not this PR's model-selection code. The PR merge-tree artifact materialized successfully, the replay guard passed, and the changed-file syntax gate passed (5 checked, 2 skipped, 0 failed). The coverage-evidence job then failed before PR-controlled tests while building the trusted coverage tool image: Could not materialize base Python locks: trusted uv executable reported an unexpected version or exit status. This is owned by the central .github coverage/toolchain layer, which is read-only in this fast-mlsirm writer scope. I am therefore not changing product code, weakening the gate, rerunning the central workflow from this repository, or dismissing the formal review. Re-evaluate this unchanged exact head once the trusted coverage infrastructure can produce same-head evidence.

Copy link
Copy Markdown
Contributor Author

@opencode-agent Please re-review the unchanged current exact head 42b849c4c05c4464205c3afe71e39ce8c0cdae2a. The current CHANGES_REQUESTED is solely the central coverage-evidence failure from run 32124705657; the owning ContextualWisdomLab/.github main has since advanced with the flat-lock/relative-include coverage-tooling repair. Re-evaluate this exact head only, preserve the relation-governance/Rust-ownership scope, and do not transfer the prior failed central evidence.

Copy link
Copy Markdown
Contributor Author

@opencode-agent Please re-review unchanged exact head 42b849c4c05c4464205c3afe71e39ce8c0cdae2a against the current central coverage implementation. Repository CI, Security Scan, CodeQL, and Semgrep are terminal-success on this SHA. The existing formal CHANGES_REQUESTED maps to central run 32124705657, which predates .github main b71a02a310e77f70c1e59f4719f6857cb33ca886 and its trusted-uv/flat-lock correction. Please generate fresh same-head formal evidence rather than carrying forward the superseded central tooling failure.

Copy link
Copy Markdown
Contributor Author

@opencode-agent Please re-review unchanged exact head 42b849c4c05c4464205c3afe71e39ce8c0cdae2a with the current central workflow. Repository CI, Security Scan, CodeQL, and Semgrep are terminal-success; the only formal REQUEST_CHANGES is prior coverage-evidence infrastructure failure. Reassess the relation-aware structural-selection governor, its separation from factor retention/Rust numerical work, and current same-head coverage evidence without transferring that superseded tooling verdict.

Copy link
Copy Markdown
Contributor Author

Superseded by #1008 at identical source SHA 42b849c4c05c4464205c3afe71e39ce8c0cdae2a. Repository CI/Security/CodeQL/Semgrep are terminal-success and no unresolved review threads remain. The formal REQUEST_CHANGES is solely central coverage-evidence run 32124705657; #1008 supplies a fresh current-workflow PR event without source churn, review dismissal, gate weakening, or force-push. Closing this predecessor unmerged avoids duplicate landing vehicles.

@seonghobae seonghobae closed this Aug 19, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant