Home/Research

Codon-level equivariance: what it fixes, and what it does not

The genetic code is degenerate, and most mRNA models learn that the expensive way. Equi-mRNA encodes it in the representation instead. The mechanism, the numbers, and the thing people get wrong about synonymous codons.

RResearch teamDeepBio ScientificAug 14, 2026 · 2 min read

Most amino acids are encoded by more than one codon. GCU, GCC, GCA and GCG all specify alanine. A model that treats an mRNA sequence as ordinary text has to discover that relationship from data, separately, for every motif in which it ever appears. That is expensive, and it makes behaviour on an unfamiliar spelling of a familiar motif hard to predict.

The symmetry is known, so give it to the model

Prior codon-level models tokenize by codon but leave the synonymous relationship implicit. Equi-mRNA encodes it directly: synonymous codons are mapped onto cyclic subgroups of SO(2), so a synonymous substitution corresponds to a rotation the representation is equivariant to by construction. The prior is held in place two ways, an auxiliary equivariance loss during training and symmetry-aware pooling when codon-level representations are aggregated to the sequence level.

The bet: geometry the architecture already knows is geometry the model does not have to spend parameters memorising. What is left over is capacity for the biology we cannot yet write down.

What the numbers say

Trained on 25M sequences drawn from 56M RefSeq entries and evaluated across six biological benchmarks, the paper reports, relative to its vanilla baselines:

property-prediction accuracy:  up to ~10% better
sequence realism:              up to ~4.3x better (Frechet BioDistance)
functional preservation:       ~28% better (generated sequences)
parameters:                    15M, vs 50M-82M for the compared baselines

Read those carefully. They are computational benchmark results on the evaluated tasks, against the baselines the paper names. "Up to" means the best task, not the average. None of them is a measured therapeutic outcome, and we would rather say so here than let the numbers travel without the qualifier.

The parameter count is not a vanity metric. It is a proxy for how much of a baseline is spent relearning a symmetry the encoding gives Equi-mRNA for free.

The part people get wrong

Equivariance here does not mean synonymous codons are interchangeable. They are not. Codon choice affects expression, folding kinetics and stability, which is exactly why the relationship is worth modelling explicitly rather than assuming away. What the model encodes is that these codons preserve amino-acid identity, not that they preserve behaviour.

One sign the representation tracks real biology is what falls out of it unsupervised: the learned codon rotations correlate with GC-content bias (r = 0.98) and with tRNA gene-copy abundance (ρ = -0.69), two independent, well-characterised signals the model was never trained to predict. That is evidence of biological interpretability. It is not therapeutic validation, and we do not report it as such.

The architecture, the per-benchmark numbers against named baselines, and the limitations are on the model card.

Keep reading

All posts →
Product

Early access to Helixir opens to research teams

Sep 11, 2026 / 1 min read
Research

MIRAGE: a benchmark for protein-family generalization in affinity prediction

Sep 6, 2026 / 1 min read
Research

A random forest beat two frontier co-folders. That is the benchmark talking.

Sep 6, 2026 / 5 min read