Repository navigation
Make EVA reproducibility easier to run and verify - #20
Merged
Merged
Conversation
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
The repository now provides a concise README and separate CPU result-recalculation, GPU inference, training, CLI and resource guides. Six reviewer software requirements are mapped to public commands and dated evidence, with missing historical resources and unreleased code/DOI work stated explicitly.
Milena completes normally with exit code 0 when inference, metrics and plots succeed.
--strict-referencepreserves exit code 2 for reference mismatch. The original inputs, weights, scoring protocol and numerical results are unchanged. CPU configuration checks also work without eagerly importing the fine-tuning GPU backend.Validation: GitHub CPU checks passed all steps on both Python versions, including 139 tests per version with no skips, 149 local links, CFF/input checks and independent wheel imports. GPU default and strict runs each scored all 135 Milena inputs on one A100 with batch size 1; predictions exactly matched the saved fresh vector, returning 0 and 2 respectively. After the import fix, synthetic pretraining, midtraining and fine-tuning each completed two steps and exact checkpoint round trips. README rendering and diagnostic plots were inspected.
The measured Spearman remains 0.8394237924835843 versus reference 0.8360456283218484. Formal release, code DOI, full comparison-method reproduction and unresolved paper resources remain separate work.