OIL, the Omics Integrated Library, was presented at the MIT Molecular Machine Learning Conference 2026 by Ivan Garibay, Mehdi Yazdani-Jahromi, Ozlem Ozmen Garibay, Sina Abdidizaji and Ali Khodabandeh Yalabadi.
It is an open framework assembling a catalogue of 2.39 billion records and 2.349 trillion nucleotides across fifteen sources, spanning metagenomics, eukaryotes, viruses, pathogens, regulatory RNAs, antibodies and patents. Coding sequence is kept in canonical reading frame at codon resolution, intergenic sequence is kept as a separate stream, and every record carries provenance. Only observed nucleotides are admitted: nothing is back-translated from protein, and 99.98% of the deduplicated catalogue is native cDNA.
Upstream licences prevent redistributing the assembled sequences, so the release is the framework that rebuilds them. We will release all the scripts and the automated procedure later this year.
Read the paper, or the write-up on why we built it.


