Home/Research

OIL: an open codon-resolved nucleotide framework, presented at MoML 2026

The Omics Integrated Library assembles 2.35 trillion nucleotides across fifteen sources, codon-resolved and with provenance on every record. Scripts and the automated procedure follow later this year.

DDeepBio ScientificDeepBio ScientificSep 1, 2026 · 1 min read

OIL, the Omics Integrated Library, was presented at the MIT Molecular Machine Learning Conference 2026 by Ivan Garibay, Mehdi Yazdani-Jahromi, Ozlem Ozmen Garibay, Sina Abdidizaji and Ali Khodabandeh Yalabadi.

It is an open framework assembling a catalogue of 2.39 billion records and 2.349 trillion nucleotides across fifteen sources, spanning metagenomics, eukaryotes, viruses, pathogens, regulatory RNAs, antibodies and patents. Coding sequence is kept in canonical reading frame at codon resolution, intergenic sequence is kept as a separate stream, and every record carries provenance. Only observed nucleotides are admitted: nothing is back-translated from protein, and 99.98% of the deduplicated catalogue is native cDNA.

Upstream licences prevent redistributing the assembled sequences, so the release is the framework that rebuilds them. We will release all the scripts and the automated procedure later this year.

Read the paper, or the write-up on why we built it.

Keep reading

All posts →
Product

Early access to Helixir opens to research teams

Sep 11, 2026 / 1 min read
Research

MIRAGE: a benchmark for protein-family generalization in affinity prediction

Sep 6, 2026 / 1 min read
Research

A random forest beat two frontier co-folders. That is the benchmark talking.

Sep 6, 2026 / 5 min read