Evo 2 40B
Release in the Evo 2 family · version evo2_40b
Maintained by Arc Institute14
The 40-billion-parameter Evo 2 checkpoint, with 50 layers and a context length of up to 1 million base pairs. It was trained autoregressively on OpenGenome2 on 2,048 GPUs, per the Savanna README; a separate evo2_40b_base checkpoint trained at 8,192-token context is also published.146
- Model hub: Model card (Hugging Face) (external site: huggingface.co)
- Repository: GitHub repository (external site: github.com)
- Repository: Savanna training framework (external site: github.com)
- Dataset hub: OpenGenome2 dataset (external site: huggingface.co)
- Paper: Nature paper (doi:10.1038/s41586-026-10176-5) (external site: doi.org)
Availability and license
Overall availability
Downloadable from Hugging Face without gating.12
Availability is separate from permission: read the license before using or redistributing.
Apache License 2.0 (evo2 inference code) (external site: raw.githubusercontent.com)5
Apache License 2.0 (Savanna training code) (external site: raw.githubusercontent.com)7
The OpenGenome2 training dataset is also published under Apache 2.0 on Hugging Face. The evo2 and Savanna LICENSE files also carry the notices for bundled third-party code, including a BSD-style NVIDIA license and the MIT license for Fairseq code. The Savanna training framework is hosted in a personal GitHub account and describes itself as maintained by a small team and not production-ready.11576
Model-disclosure tier
Open-weight, plus published inference code, training code, and training recipe, and at least documented training-data composition.
Public materials checklist
| Item | Status | Notes and evidence |
|---|---|---|
| WeightsCan the general public download the model parameters for this release? | Public | Weights are on Hugging Face without gating; Savanna-format checkpoints (savanna_evo2_40b) are also published.123 |
| Inference codeIs code for running the model published? | Public | The evo2 package (built on the Vortex inference code) is on GitHub and PyPI under Apache 2.0.45 |
| Training codeIs the code used to train the model published? | Public | The README states Evo 2 was trained with Savanna, an Apache-2.0 pretraining framework whose README lists Evo 2 40B among the models trained with it.467 |
| Training-data informationPublic = the training data itself can be obtained. Partial = composition or sources are documented without full access. | Public | OpenGenome2 is downloadable from Hugging Face, both as raw FASTA files and as the preprocessed JSONL sequences used for pretraining.114 |
| Training recipeAre the training configuration and procedure documented in enough detail to follow? | Public | The Nature paper's Supplementary Information documents the two training phases (pretraining at 1,024 and then 8,192 tokens, then multi-stage context extension to 1 million tokens) with learning rates, batch sizes, iteration and token counts, optimizer settings, and the context-extension schedule. Savanna publishes 40B pretraining, context-extension, and data configuration files, and the OpenGenome2 card lists per-source data weights for each phase.910811 |
| Evaluation materialsPublic = evaluation code or prompts that let others re-run the evaluations are published. Partial = results only. | Partial | The repository includes notebooks for zero-shot BRCA1 variant effect prediction and an exon classifier, but it was not confirmed that all reported evaluations can be re-run.4 |
What it is useful for
Likelihood scoring of DNA sequences (such as zero-shot variant effect prediction), embeddings for downstream models, and DNA sequence generation.4
Run and use notes
- The README states that the 40B model needs FP8 through NVIDIA Transformer Engine and an NVIDIA Hopper GPU for numerical accuracy, and that it needs multiple H100 GPUs; Vortex splits the model across the available CUDA devices.4
- Documented requirements are Linux (WSL2 with limited support), CUDA 12.1 or later, cuDNN 9.3 or later, and Python 3.11 or 3.12.4
Organization context
Provenance and derivatives
Trained from scratch by Arc Institute and collaborators on OpenGenome2 using Savanna. OpenGenome2 draws on public sources including GTDB, IMG/PR, IMG/VR, metagenomic data, NCBI, Ensembl, EPDnew, RNAcentral, and Rfam.411
- Derived from: OpenGenome2 (external site: huggingface.co) — Pretraining dataset published by Arc Institute.
Catalog records that name this entry in their provenance:
Other releases in the Evo 2 family
- Evo 2 20BModel-disclosure tier (USASI rubric v0.1): Open-weight
U.S. eligibility
Eligible · basis: U.S.-governed project
The checkpoint is published in Arc Institute's Hugging Face organization and documented in Arc's GitHub repository. Arc Institute is a Palo Alto, California nonprofit research organization listed by the IRS as a 501(c)(3). Arc describes Evo 2 as developed by Arc and NVIDIA scientists with collaborators at several universities and at Goodfire.14131412
Sources
This listing is not an endorsement, a safety assessment, or a federal approval.