Skip to content
GENTEL-labPublic

About

RNA-life creation is coming ~(EVA2)

Resources

Contributing

Stars

67 stars

Watchers

8 watching

Forks

Repository files navigation

EVA — RNA foundation model

EVA

A Long-Context Generative Foundation Model for Versatile RNA Design

Paper on bioRxiv Models on Hugging Face OpenRNA dataset CPU checks on main Apache 2.0 license

Get started   ·   Reproduce the paper   ·   Documentation   ·   Website   ·   Cite EVA

EVA brings RNA sequence scoring and conditional sequence modeling into one framework, with public model code, checkpoints and selected paper reproduction workflows. Model weights and large datasets are downloaded separately.

1.4B-parameter MoE flagship model; 8,192-token context window; trained on OpenRNA v1

Get started

01   Explore on CPU

Recalculate a bundled benchmark result from saved predictions.

Python 3.10 / 3.11 · No model download

Try the quick start →

02   Run on GPU

Score all 135 Milena sequences from a pinned EVA checkpoint.

Linux / NVIDIA GPU · Docker runtime

Run the benchmark →

03   Train & fine-tune

Start with a small training and checkpoint save/reload workflow.

Source checkout · Compatible GPU runtime

Explore training →

Quick start

A first result with Python's standard library:

git clone https://github.com/GENTEL-Lab/EVA.git
cd EVA
python3 scripts/reproduce_historical_benchmark.py \
  --artifact-only --output results/milena_stored

Expected: 135 stored predictions, Spearman 0.84, exit code 0. The output directory contains report.json and sequence_label_audit.csv. This recalculates an archived result; the GPU path below runs new inference. Use a new output directory when rerunning.

Reproduce the paper

The representative workflow connects pinned checkpoint → 135 input sequences → new predictions → metrics & plots. Inputs, model files and the scoring protocol are bound to version and checksum records.

Run the GPU example · Environment setup and complete inference command

From the repository root, build and enter the runtime on the host:

mkdir -p checkpoint results
docker build -f scripts/docker/Dockerfile -t eva:local .
docker run --rm -it --gpus device=0 --name eva-repro \
  -v "$PWD":/eva -w /eva eva:local bash

Then run inside the container:

python scripts/download_reproduction_checkpoint.py \
  --model EVA_1.4B_CLM --destination checkpoint
python scripts/reproduce_milena.py \
  --checkpoint checkpoint/EVA_1.4B_CLM --output results/milena_fresh

See Installation for source installs, runtime versions and Singularity / Apptainer instructions.

Expected results & measured runtime · Reference comparison, exit codes and hardware

Spearman correlations in this guide are rounded to two decimal places.

Measurement Recorded result
New-inference Spearman 0.84
Archived reference 0.84

The full-precision values are 0.8394237924835843 (new inference) and 0.8360456283218484 (archived reference), a difference of 0.0033781641617359748. Both round to 0.84; the prediction vectors and full-precision correlations differ.

Runtime measurement Recorded result
Scoring time About 12 seconds for all 135 sequences
Peak PyTorch-allocated GPU memory 2.94 GiB
Environment One NVIDIA A100-SXM4-80GB; batch size 1; pinned Docker runtime

Default mode returns 0 when inference, metrics and plots complete successfully. Add --strict-reference to return 2 when comparison at the stored reference precision fails. Input, dependency, scoring and plotting failures remain errors in either mode.

These are measurements for this example, not minimum hardware requirements or estimates for other workloads. Downloads, loading, hashing and plotting add overhead. A successful example does not establish reproduction of every paper experiment. See dated validation records.

Protocol & outputs   ·   Paper workflow coverage   ·   Validation records

Documentation

GuideWhat you will find
Using EVA → Scoring and generation, conditioning, batch configuration, and input/output formats.
Reproduction → CPU recalculation, GPU inference, training, and the scope of verified paper workflows.
Models & data → EVA checkpoints, OpenRNA, deposited analysis data, and historical comparison-model resources.
Development → Installation for contributors, tests, pull requests, and reproducible bug reports.

Repository map · Where the code and reproduction assets live
Directory Contents
eva/ Model and tokenizer
tools/ Scoring, generation and evolution CLIs
training/ Pretraining, midtraining, fine-tuning and evaluation
examples/ Notebooks, paper benchmarks, sample data and configurations
scripts/ Reproduction commands, checks and container recipes
docs/ Guides and illustrations
tests/ Regression tests

Start with examples to choose a tutorial or benchmark. The wheel contains the importable model and CLI packages; use the complete source checkout for training, notebooks and paper-reproduction inputs.

Citation

Cite the paper and the exact Git commit used in your work. Software citation metadata are available in CITATION.cff.

The v1.2.1 source archive is available from the versioned GitHub release and Figshare code archive. The archive metadata identifies the exact source commit and SHA256 checksums. Record the model and dataset revisions listed in artifact provenance. See validation & release status for the validation records.


Apache-2.0 license   ·   Contribute   ·   Report an issue   ·   EVA website

External datasets and third-party checkpoints retain their upstream terms.

About

RNA-life creation is coming ~(EVA2)

Resources

Contributing

Stars

67 stars

Watchers

8 watching

Forks

Releases

Packages

Contributors

Languages