What it is
Evo 2 is a DNA foundation model trained on 9 trillion nucleotides spanning all domains of life, with a 1-million-token context at single-base resolution. Without task-specific fine-tuning it predicts the functional impact of genetic variation, from noncoding pathogenic mutations to clinically significant BRCA1 variants. It also generates coherent genome-scale sequences and, probed internally, has learned features like exon-intron boundaries and transcription-factor binding sites.
Why it matters
At 40 billion parameters and a megabase context, Evo 2 is the largest and most general biological sequence model to date, unifying prediction and design across the tree of life in one system. The team released the model weights, code, and the OpenGenome2 dataset, so the wider field can build on it.
Underlined numbers link to their source. Every metric and quoted figure is listed under Sources and data below.
Filed undergenomics, foundation models, DNA, variant prediction, Evo 2
Watch
A short explainer of this result.