Skip to content

Make EVA reproducibility easier to run and verify - #20

Merged
kevinhyj merged 2 commits into
mainfrom
repository-usability-20260909
Sep 9, 2026
Merged

kevinhyj merged 2 commits into
mainfrom
repository-usability-20260909

Conversation

@kevinhyj

@kevinhyj kevinhyj commented Sep 9, 2026 •

Copy link
Copy Markdown
Collaborator

The repository now provides a concise README and separate CPU result-recalculation, GPU inference, training, CLI and resource guides. Six reviewer software requirements are mapped to public commands and dated evidence, with missing historical resources and unreleased code/DOI work stated explicitly.

Milena completes normally with exit code 0 when inference, metrics and plots succeed. --strict-reference preserves exit code 2 for reference mismatch. The original inputs, weights, scoring protocol and numerical results are unchanged. CPU configuration checks also work without eagerly importing the fine-tuning GPU backend.

  • Add Python 3.10/3.11 CI for regression tests, documentation/CFF/input checks, source-archive execution, package builds, wheel installation outside the checkout and installed CLIs. External resource access is reported separately.
  • Recover and hash 15 files from three historical RNA model snapshots, plus 13 additional cached comparison-model assets. This does not imply public availability, full runtime recovery or new competitor inference.
  • Add contribution guidance and issue/PR templates; preserve model architecture and existing directories.

Validation: GitHub CPU checks passed all steps on both Python versions, including 139 tests per version with no skips, 149 local links, CFF/input checks and independent wheel imports. GPU default and strict runs each scored all 135 Milena inputs on one A100 with batch size 1; predictions exactly matched the saved fresh vector, returning 0 and 2 respectively. After the import fix, synthetic pretraining, midtraining and fine-tuning each completed two steps and exact checkpoint round trips. README rendering and diagnostic plots were inspected.

The measured Spearman remains 0.8394237924835843 versus reference 0.8360456283218484. Formal release, code DOI, full comparison-method reproduction and unresolved paper resources remain separate work.

@kevinhyj
kevinhyj merged commit 211db94 into main Sep 9, 2026
2 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant