TL;DR
Illumina launched SpliceAI2 today, the successor to its 2019 SpliceAI model that predicts how DNA variants alter RNA splicing — trained on a dataset more than 100 times larger than the original. In a 7,504-person rare-disease cohort it found 17% more disease-relevant variants than any other splicing model, and it beat SpliceAI, Pangolin, and Google DeepMind’s AlphaGenome across all three benchmarks. The code, trained weights, and precomputed predictions for 4 billion possible variants are open for academic and non-commercial research.

What SpliceAI2 actually does
SpliceAI2 comes out of Illumina’s BioInsight AI Lab and predicts how genetic variants alter RNA splicing. The original SpliceAI answered one question — would a cell splice at a given location. SpliceAI2 answers three: which positions in a gene are used as splice sites and how frequently, which splice sites connect through splice junctions, and which full-length RNA transcript isoforms are produced.
The practical twist: it needs only a DNA sequence as input, so researchers get transcript-level analysis without RNA data from hard-to-obtain tissues. The model processes 196,608 base pairs of genomic sequence with roughly 13 million trainable parameters, and it was trained end-to-end on 314,745 RNA sequencing samples from humans and nine other mammalian species — covering more than 46 million observed splice junctions — plus 330 long-read samples from the public ENCODE project.
The numbers that matter
Across three independent benchmarks, SpliceAI2 beat the original SpliceAI, Pangolin, and Google DeepMind’s AlphaGenome — with the AlphaGenome comparisons run independently by collaborators at the University of Oxford. On GTEx cryptic splice variant detection it scored an auPRC of 0.77 against 0.66 for the next best model; on splice site usage quantification it reached a Spearman correlation of 0.63 against 0.47, which Illumina describes as a 34% improvement in splice-site usage quantification versus the next best model.
The rare-disease test is the one clinicians will care about. Across 7,504 probands from the Genomics England 100,000 Genomes Project, variants prioritized by SpliceAI2 identified 17% more disease-associated variants than any other tested splicing model at matched confidence thresholds. Against its own predecessor it found 33% more disease-relevant splice variants at a 2X confidence interval and 66% more at a 4X interval.
Population-scale validation drew on more than 627,000 genomes from gnomAD, TOPMed, and UK Biobank: the variants SpliceAI2 scored highest were strongly depleted, close to the depletion seen for protein-truncating loss-of-function mutations. And in UK Biobank proteomic data from 36,764 people, carriers of higher-scoring variants had lower plasma protein levels — a Pearson correlation of -0.50, the strongest of any model tested.
Why it finds variants the others miss
Roughly half the cryptic splice variants SpliceAI2 identified sit deep inside intronic regions — and variants more than 50 base pairs into an intron are typically missed by exome sequencing, which is what most clinical testing still uses. Predicted splice-altering variants accounted for 15% of the excess genetic burden in the rare-disease cohort. That is the genome’s dark matter: the 98% of DNA that doesn’t code for protein but can still cause disease when splicing goes wrong.
The model also carries a fine-tuning framework that conditions predictions on the expression of 147 RNA binding proteins, capturing tissue-specific splicing differences across 48 GTEx tissues. The authors report the model independently learned sequence motifs recognized by real splicing regulators without being explicitly taught them, and that the framework was adapted to disease states including SF3B1-mutant tumors and myotonic dystrophy.
Kyle Farh, vice president of the BioInsight AI Lab, put the ambition in one line: Illumina is advancing AI “to systematically shrink the portion of the genome that remains uninterpretable.”
How to use it
The bar is high: the original SpliceAI from 2019 was cited in more than 3,400 publications and is incorporated into splice variant interpretation recommendations from ClinGen, the clinical research body that sets standards for clinical genomics. SpliceAI2 ships through Illumina’s BioInsight Platform applications, including DRAGEN Annotation, Emedgene, and Illumina Connected Insights.
For researchers, the source code, trained model weights, and precomputed predictions live on GitHub and Hugging Face for academic and non-commercial research use — all possible single-nucleotide variants within human gene bodies (4 billion) plus population indels (150 million). It installs through PyPI and needs a CUDA-capable GPU. The summary score runs 0 to 1, with recommended thresholds of 0.1 for high recall, 0.25 for balanced precision and recall, and 0.5 for high precision — the SpliceAI equivalents of 0.2, 0.5, and 0.8.
And it is part of a suite: SpliceAI2 joins PromoterAI and PrimateAI-3D, which Illumina says together let researchers identify up to twice as many variants with predicted biological impact. Rare-disease research, hereditary cancer testing research, and drug discovery are the target applications — anywhere variants of uncertain significance are the bottleneck.