← All models
Model · Foundation model

Equi-mRNA

A codon-level equivariant language model for mRNA: the representation underneath expression prediction, stability assessment, generation, and therapeutics design in the co-scientist.

NeurIPS 2025 · Peer-reviewed · arXiv 2508.15103 · Yazdani-Jahromi, Khodabandeh Yalabadi & Ozmen Garibay
Overview

The genetic code is degenerate: most amino acids are encoded by more than one codon. GCU, GCC, GCA, and GCG all specify alanine: different codons encoding the same amino acid. (Synonymous substitutions are not inert: they can still affect expression and stability, which is part of why the symmetry is worth modelling explicitly.) A model that treats mRNA as ordinary text has to learn that equivalence from data, separately, for every motif it ever appears in.

Equi-mRNA encodes the structure directly instead. Synonymous codons are mapped to cyclic subgroups of SO(2), so codon substitutions that preserve amino-acid identity become rotations the model is equivariant to by construction. Geometry the architecture already knows is geometry it doesn't have to spend parameters memorizing.

Architecture
  • Codon-level equivariant encoder: 15M parameters (GPT-2 backbone), 5M-parameter Mamba-hybrid variant also trained
  • Symmetry group: cyclic subgroups of SO(2), mapped from synonymous-codon degeneracy
  • An auxiliary equivariance loss enforces the prior during training
  • Symmetry-aware pooling (polar, DFT-based, or SO(2) mean) aggregates codon-level representations into a sequence-level one
  • Rotation basis optionally learned on the Stiefel manifold St(2,d) rather than fixed to a standard 2D plane
Training & data

25M coding sequences drawn from 56M RefSeq entries (20–512 codons, canonical bases only), pretrained on 32 NVIDIA H100 GPUs. Evaluated across six biologically-driven benchmarks, each split 70/15/15 train/validation/test:

DatasetTaskSize
MLOSFlu-vaccine mRNA expressionSanofi-Aventis influenza vaccine mRNAs, ~1,700 nt543 sequences
mRFPFluorescent-protein expressionSynthetic constructs, 678 nt fixed length1,459 sequences
E. coliProtein expression (high/low)171–3,000 nt6,348 sequences
Tc-RiboswitchRegulatory switching factor67–73 nt355 sequences
iCodonHuman mRNA stability30–1,497 nt41,123 sequences
SARS-CoV-2 DegradationVaccine degradation rate107 nt fixed length2,400 sequences
Benchmarks

Against baselines several times its size, Equi-mRNA holds its own. All figures below are the paper's reported results on its evaluated tasks, relative to its vanilla baselines: computational benchmarks, not measured therapeutic outcomes.

up to ~10%
Property-prediction accuracy, vs. the paper's vanilla baselines
4.3×
Generative fidelity, Fréchet BioDistance
~28%
Better functional-property preservation in generated sequences
15M
Params, vs. 50–82M for compared baselines

Score vs. parameter count

One scatter per dataset: model size (log scale) against score. Only models Table 2 reports a parameter count for can appear here; RNA-FM, CALM, and mRNA-FM don't state one.

1.00.80.60.40.20.010M100M1BParameters (log scale)
RNABERTAido mRNACodonBERTEqui-mRNA (5M)Equi-mRNA (15M)

Ranked by task

Equi-mRNA (15M)0.855
Equi-mRNA (5M)0.853
HELM (CLM)0.849
CodonBERT0.832
HELM (MLM)0.822
GPT-2 (CLM)0.815
RNA-FM0.800
GPT-2 (MLM)0.753
Aido mRNA0.683
mRNA-FM0.564
CALM0.546
RNABERT0.400
Ablation

Table 1 of the paper: the top-performing configuration from each rotation strategy, on the 1M-sequence ablation subset. Fuzzy θ (soft, learned codon-rotation distributions with a Stiefel basis and the equivariance loss) was carried forward to full-scale training as Equi-mRNA.

VariantStiefelEquiv. lossE. coliMLOSTc-Ribo.mRFPCOV Deg
Vanilla––0.5800.633 ± .140.6980.7970.779
Fixed θ✓–0.6020.667 ± .190.6880.8400.803
Learned θ✓–0.6330.657 ± .100.7010.8710.790
Fuzzy θ✓✓0.6050.691 ± .140.7360.8440.820
Vanilla
E. coli0.580
MLOS0.633 ± .14
Tc-Ribo.0.698
mRFP0.797
COV Deg0.779
Fixed θStiefel
E. coli0.602
MLOS0.667 ± .19
Tc-Ribo.0.688
mRFP0.840
COV Deg0.803
Learned θStiefel
E. coli0.633
MLOS0.657 ± .10
Tc-Ribo.0.701
mRFP0.871
COV Deg0.790
Fuzzy θStiefelEquiv. loss
E. coli0.605
MLOS0.691 ± .14
Tc-Ribo.0.736
mRFP0.844
COV Deg0.820
Generative fidelity

Sequences generated by autoregressive codon sampling, scored against true suffixes by Fréchet BioDistance (lower is better) across sampling temperatures. Vanilla GPT-2 stays flat and unstable; both Equi-mRNA variants decline sharply as temperature rises, with the 5M model reaching ≈76 at T=1.0, a ≈33× gap from vanilla at full training scale. The paper's own headlined ≈4.3× gain comes from the equivariant prior alone, in the controlled 1M-sequence ablation.

250020001500100050000.20.40.60.81.0Sampling temperatureFBD (↓ better)Vanilla GPT-2 · T=0.2 · FBD 2561.81Vanilla GPT-2 · T=0.4 · FBD 2561.99Vanilla GPT-2 · T=0.6 · FBD 2562.24Vanilla GPT-2 · T=0.8 · FBD 2562.67Vanilla GPT-2 · T=1 · FBD 2562.78Equi-mRNA (5M) · T=0.2 · FBD 819.18Equi-mRNA (5M) · T=0.4 · FBD 579.21Equi-mRNA (5M) · T=0.6 · FBD 341.68Equi-mRNA (5M) · T=0.8 · FBD 129.43Equi-mRNA (5M) · T=1 · FBD 76.13Equi-mRNA (15M) · T=0.2 · FBD 802.89Equi-mRNA (15M) · T=0.4 · FBD 310.4Equi-mRNA (15M) · T=0.6 · FBD 196.15Equi-mRNA (15M) · T=0.8 · FBD 172.02Equi-mRNA (15M) · T=1 · FBD 177.77
Vanilla GPT-2Equi-mRNA (5M)Equi-mRNA (15M)
Biological interpretability

One sign the representation tracks real biology, rather than a benchmark artifact, is what falls out of it unsupervised, fine-tuned on the human coding transcriptome (GRCh38). The learned codon rotations correlate with GC-content bias (Pearson r = 0.98, R² = 0.97, p < 10⁻¹¹) and with tRNA gene-copy abundance (Spearman ρ = −0.69, p < 10⁻⁶): two independent, well-characterized biological signals the model was never directly trained to predict. These associations are evidence of biological interpretability; they do not establish therapeutic validation.

Publication

Mehdi Yazdani-Jahromi, Ali Khodabandeh Yalabadi, Ozlem Ozmen Garibay, University of Central Florida.

Mehdi Yazdani-Jahromi is a DeepBio Scientific co-founder. All three authors list the University of Central Florida on the publication; this note does not make a separate ownership or licence claim.

Peer-reviewed at NeurIPS 2025 (39th Conference on Neural Information Processing Systems). The full paper (architecture, training data, and evaluation) is on arXiv.

arXiv 2508.15103 ↗
Used in

The representation behind four programs in the co-scientist:

  • Expression prediction
  • Stability assessment
  • mRNA generation
  • Therapeutics design
See it running in the platform →

FAQ

A codon-level language model for mRNA that encodes synonymous-codon symmetry directly into its architecture, as cyclic subgroups of SO(2), rather than leaving the model to infer that structure from data.

Read the primary source.

The paper covers the full architecture, training procedure, and evaluation methodology.

Read on arXiv ↗