Home/Research

Codon resolution at metagenomic scale: why we built OIL

Codon corpora stop at reference organisms. Metagenomic corpora throw the codons away by storing coding sequence as protein. OIL is 2.35 trillion nucleotides that keep both, with provenance on every record and nothing back-translated.

RResearch teamDeepBio ScientificSep 1, 2026 · 5 min read

If you are building a language model over protein-coding sequence, you have to answer a question most papers leave implicit: at what resolution do you represent the sequence? Amino acids, raw nucleotides, or codons?

For mRNA therapeutics the answer is not close. Design largely is the choice among synonymous codons, and that choice sets translational efficiency, transcript stability and immunogenicity. A model that cannot see codons cannot see the design space.

So the corpus has to keep them. None of the open ones does, at least not together with the biological diversity that makes codon usage worth modelling in the first place.

The gap, stated precisely

Frame-independent DNA corpora like OpenGenome2 and GenomeOcean keep genes inside unfiltered contigs, under six-frame and strand uncertainty, with tokenization that crosses codon boundaries. Metagenomic corpora like OMG store coding sequence as translated amino acids, from which synonymous codons cannot be recovered at all. Codon-aware corpora do retain codons, but every one of them is built from curated reference genomes and carries no non-coding sequence.

That leaves a hole exactly where the signal is. Codon preference reflects GC content, tRNA pools and translational selection, and those vary enormously across life, most of which is uncultivated. The widest natural variation in codon usage is metagenomic, and the metagenomic corpora are precisely the ones that discard it.

What OIL is

The Omics Integrated Library assembles a catalogue of 2.39 billion records, 2.349 trillion nucleotides, from fifteen sources spanning metagenomics, eukaryotes, viruses, pathogens, regulatory RNAs, antibodies and patents. Every coding sequence is stored in canonical reading frame at codon resolution; every intergenic sequence as nucleotides; every record with provenance.

metagenomic coding        1,890,800,696   79.19%
metagenomic intergenic      253,528,927   10.62%
eukaryotic reference        172,666,812    7.23%
plant and fungal             43,056,471    1.80%
viral and pathogen           27,101,444    1.14%
antibody and patent             377,528    0.02%

After cross-source deduplication the coding catalogue holds 289,485,932 representative sequences, roughly 6.7 times the 43.5M of SynCodonLM, the largest openly released codon corpus before it. The intergenic stream is kept separately, 1.77 billion staged fragments reduced to 254 million representatives, so genomic context is available to a model without being imposed on it.

Only observed nucleotides

The inclusion policy is the part we would defend hardest. Every sequence enters as observed nucleotides. Nothing is back-translated from protein, and nothing is synthesised. Across the deduplicated catalogue, 99.98% of records are native cDNA, and every patent entry carrying a codon-optimised flag is catalogue-only, excluded from default configurations.

The reason is narrow and important. Back-translation manufactures exactly the statistics a codon model is supposed to read. Synonymous usage encodes GC content, species-specific bias and functional motifs; an optimiser invents those patterns rather than recording them. A corpus that mixes the two teaches a model to recognise an algorithm's preferences and call it biology.

Recovery was attempted for 52,056 amino-acid-only candidates and succeeded for 15,721. The rest stayed out. A smaller corpus you can trust beats a larger one you cannot.

Deduplication, judged against biology

Deduplication runs in two stages: MMseqs2 linclust at 80% nucleotide identity, then a semantic stage over Evo 2 embeddings. The second stage is where a corpus like this can quietly destroy its own reason for existing, because the diversity OIL exists to preserve looks, to a nearest-neighbour metric, a lot like redundancy.

So we tested the threshold where the loss would happen, on metagenomic pairs. At the source method's 0.90 threshold, the pairs being absorbed are statistically indistinguishable from a control drawn just below the cut: 0.85% versus 0.90% reach 80% nucleotide identity. In other words, at the conventional setting the filter isolates no duplicate population at all among metagenomic sequence; it is removing diversity. Precision rises sharply above it: 63.3% of absorbed pairs reach 80% identity at cosine 0.95, and 94.3% at 0.97.

We use 0.95, which is more conservative than the method we adapted, chosen where duplication becomes evident rather than where convention put it. The embedding space is also strongly anisotropic, mean cosine similarity between random pairs is 0.861, so a whitening transform is fitted once and frozen in the release manifest with its checksum.

A catalogue, not a build

Each sequence is stored once in a single relational table with its source, molecular type, quality and cluster membership. A training set is then a declarative query over that table, not a separate artifact. OIL-Earth, the broad pretraining configuration, is a predicate. OIL-Med, narrowed to the sources most relevant to mRNA and antibody modelling, is another.

The point is that the assumptions a fixed corpus bakes in become columns you can query and change. Want a different identity threshold? Cluster membership is retained per record, so it is a query rather than a rebuild.

Benchmark leakage is handled the same way. Every record is matched against benchmark datasets including iCodon and the Diverse Genomic Embedding Benchmark and flagged; configurations draw only on unflagged sequences, and flagged records stay in the catalogue so someone targeting a different benchmark can rescreen. The overlap is small: 7,126 coding records, and 39 intergenic hits falling on 12 distinct representatives.

What we are not claiming

No modelling claims. We have not trained a model on OIL and reported that it is better, and this paper does not argue that one would be. What we show is a property of the data: third-position GC ranges from 0.27 to 0.88 in the metagenomic portion against 0.36 to 0.65 in the RefSeq-derived eukaryotic set, so the metagenomic portion covers codon-usage space the reference organisms do not. That follows from composition, not from any result.

Four limitations are stated in full in the paper: prokaryotic bias within the metagenomic majority, absent per-record taxonomy, coding sequence that is predicted rather than observed in the unannotated sources, and validation by stratified sampling.

Release

Upstream licence terms prevent us from redistributing the assembled sequences, so what we release is the framework that rebuilds them: the ingestion and gene-calling pipeline, the catalogue schema, the quality and inclusion predicates, the deduplication configuration including the frozen whitening transform and its checksum, and the methods ledger. The gene caller is deterministic across the release series we used, so the coding stream reproduces from the pipeline alone given the same upstream snapshots.

We will release all the scripts and the automated procedure later this year. Until then, the paper carries the schema, the predicates and the per-source detail in its appendices, and we are happy to talk to anyone who wants to rebuild it sooner.

OIL was presented at the MIT Molecular Machine Learning Conference 2026, by Ivan Garibay, Mehdi Yazdani-Jahromi, Ozlem Ozmen Garibay, Sina Abdidizaji and Ali Khodabandeh Yalabadi. Read the paper. It is the corpus-side counterpart to Equi-mRNA: one asks how to represent a codon, the other asks what to show a model so the representation has something to learn from.

Keep reading

All posts →
Product

Early access to Helixir opens to research teams

Sep 11, 2026 / 1 min read
Research

MIRAGE: a benchmark for protein-family generalization in affinity prediction

Sep 6, 2026 / 1 min read
Research

A random forest beat two frontier co-folders. That is the benchmark talking.

Sep 6, 2026 / 5 min read