A Long-Context Generative Foundation Model for Versatile RNA Design
Get started · Reproduce the paper · Documentation · Website · Cite EVA
EVA brings RNA sequence scoring and conditional sequence modeling into one framework, with public model code, checkpoints and selected paper reproduction workflows. Model weights and large datasets are downloaded separately.
|
Recalculate a bundled benchmark result from saved predictions. Python 3.10 / 3.11 · No model download Try the quick start → |
Score all 135 Milena sequences from a pinned EVA checkpoint. Linux / NVIDIA GPU · Docker runtime Run the benchmark → |
Start with a small training and checkpoint save/reload workflow. Source checkout · Compatible GPU runtime Explore training → |
A first result with Python's standard library:
git clone https://github.com/GENTEL-Lab/EVA.git
cd EVA
python3 scripts/reproduce_historical_benchmark.py \
--artifact-only --output results/milena_storedExpected: 135 stored predictions, Spearman 0.84, exit code 0.
The output directory contains report.json and sequence_label_audit.csv.
This recalculates an archived result; the GPU path below runs new inference.
Use a new output directory when rerunning.
The representative workflow connects pinned checkpoint → 135 input sequences → new predictions → metrics & plots. Inputs, model files and the scoring protocol are bound to version and checksum records.
Run the GPU example · Environment setup and complete inference command
From the repository root, build and enter the runtime on the host:
mkdir -p checkpoint results
docker build -f scripts/docker/Dockerfile -t eva:local .
docker run --rm -it --gpus device=0 --name eva-repro \
-v "$PWD":/eva -w /eva eva:local bashThen run inside the container:
python scripts/download_reproduction_checkpoint.py \
--model EVA_1.4B_CLM --destination checkpoint
python scripts/reproduce_milena.py \
--checkpoint checkpoint/EVA_1.4B_CLM --output results/milena_freshSee Installation for source installs, runtime versions and Singularity / Apptainer instructions.
Expected results & measured runtime · Reference comparison, exit codes and hardware
Spearman correlations in this guide are rounded to two decimal places.
| Measurement | Recorded result |
|---|---|
| New-inference Spearman | 0.84 |
| Archived reference | 0.84 |
The full-precision values are 0.8394237924835843 (new inference) and 0.8360456283218484 (archived reference), a difference of 0.0033781641617359748. Both round to 0.84; the prediction vectors and full-precision correlations differ.
| Runtime measurement | Recorded result |
|---|---|
| Scoring time | About 12 seconds for all 135 sequences |
| Peak PyTorch-allocated GPU memory | 2.94 GiB |
| Environment | One NVIDIA A100-SXM4-80GB; batch size 1; pinned Docker runtime |
Default mode returns 0 when inference, metrics and plots complete successfully.
Add --strict-reference to return 2
when comparison at the stored reference precision fails. Input, dependency,
scoring and plotting failures remain errors in either mode.
These are measurements for this example, not minimum hardware requirements or estimates for other workloads. Downloads, loading, hashing and plotting add overhead. A successful example does not establish reproduction of every paper experiment. See dated validation records.
Protocol & outputs · Paper workflow coverage · Validation records
| Guide | What you will find |
|---|---|
| Using EVA → | Scoring and generation, conditioning, batch configuration, and input/output formats. |
| Reproduction → | CPU recalculation, GPU inference, training, and the scope of verified paper workflows. |
| Models & data → | EVA checkpoints, OpenRNA, deposited analysis data, and historical comparison-model resources. |
| Development → | Installation for contributors, tests, pull requests, and reproducible bug reports. |
Repository map · Where the code and reproduction assets live
| Directory | Contents |
|---|---|
eva/ |
Model and tokenizer |
tools/ |
Scoring, generation and evolution CLIs |
training/ |
Pretraining, midtraining, fine-tuning and evaluation |
examples/ |
Notebooks, paper benchmarks, sample data and configurations |
scripts/ |
Reproduction commands, checks and container recipes |
docs/ |
Guides and illustrations |
tests/ |
Regression tests |
Start with examples to choose a tutorial or benchmark. The wheel contains the importable model and CLI packages; use the complete source checkout for training, notebooks and paper-reproduction inputs.
Cite the paper and the exact Git commit used in your work. Software citation metadata are available in CITATION.cff.
The v1.2.1 source archive is available from the versioned GitHub release and Figshare code archive. The archive metadata identifies the exact source commit and SHA256 checksums. Record the model and dataset revisions listed in artifact provenance. See validation & release status for the validation records.
Apache-2.0 license · Contribute · Report an issue · EVA website
External datasets and third-party checkpoints retain their upstream terms.