Matched-pK Pearson r within each support bin. Co-folders rise steeply with the number of structures their family already had; family-disjoint controls stay flat.
Gm is the change in matched-pK accuracy from singleton families to the best-supported ones. The two co-folders carry a large positive gap; the family-disjoint controls and the classical empirical scorer do not. ΔG is the paired difference on the co-folder's own complexes, since prediction coverage differs by method.
| # | Method | 95% CIon Gm | |||
|---|---|---|---|---|---|
| Boltz-2Affinity co-folder | +0.62 | – | – | [+0.13, +1.02] | |
| Nesso-1Affinity co-folder | +0.45 | – | – | [+0.28, +0.61] | |
| gninaDocking / deep net | +0.07 | +0.38[+0.17, +0.59] | +0.52[−0.08, +1.09] | [−0.14, +0.27] | |
| RF-QSARFamily-disjoint control | -0.01 | +0.53[+0.26, +0.75] | +1.11[+0.49, +1.71] | [−0.17, +0.16] | |
| ligand-kNNFamily-disjoint control | -0.03 | +0.48[+0.25, +0.72] | +0.92[+0.22, +1.57] | [−0.17, +0.11] | |
| sminaDocking / deep net | -0.04 | +0.32[+0.09, +0.59] | +0.84[+0.10, +1.47] | [−0.18, +0.19] |
Table 1. Matched-pK window, MMseqs2 30% clustering, two-level bootstrap 95% CI. Both co-folder gaps exclude zero; every family-disjoint control interval spans it. Every paired ΔG excludes zero except gnina against Boltz-2. Restricting to matched samples moves the controls' gaps more negative than their full-sample values, so an unpaired comparison understates the interaction.
How protein-ligand affinity and pose accuracy depend on historical public protein-family support: the number of PDBbind structures through 2019 in the same MMseqs2 30% sequence-identity cluster. Standard benchmarks report one overall correlation, which averages over that variable. MIRAGE treats it as the independent variable and reports accuracy across the whole curve, for affinity and for pose.
No, and the paper is deliberate about that. Leakage would mean related families demonstrably crossing a claimed train/test boundary, and the proprietary affinity training sets cannot be reconstructed to prove it. The defensible claim is weaker and still decisive: accuracy scales with historical public family support after controlling for measured confounds, and matched family-disjoint models do not share that dependence. The paper calls this public family-support dependence and redundancy-driven inflation. Family density stays a proxy, not evidence of training exposure.
They are the most accurate methods available on well-supported families, and the paper says so. The problem is the reporting convention: a single overall correlation conflates interpolation within familiar families with transfer to novel ones. For a target with no close relatives in the PDB, the honest expectation is the low-support number, roughly 0.1 to 0.3. The pose arm adds the constructive half: for Boltz-2, supplying a ColabFold MSA raises novel-family pose success from 2% to 46% with redundant success unchanged, so part of that deficit is a missing input rather than an irreducible limit.
Only on novel families, and only because that regime removes the advantage the co-folders are actually using. On families with five or fewer prior structures a family-disjoint random forest scores 0.411 against Nesso-1's 0.324 (paired difference +0.087 [+0.007, +0.167], excluding zero) and 0.505 against Boltz-2's 0.308 on its 189-complex coverage (+0.197 [−0.018, +0.405], the larger point estimate but not separated from zero). Its accuracy is flat across the support curve because it never had family support to spend.
G_m is the change in matched-pK Pearson r from singleton families to well-supported ones: r(S_f ≥ 301) minus r(S_f = 1). A large positive G_m means performance increased with family support. It is estimated with a two-level bootstrap, resampling families and then ligands within families, so the interval reflects both sources of variability.
They measure different stringencies, and both are reported. A 30% sequence split is a far weaker generalization test than singleton-family stratification: a ligand-only model reaches 0.440 on LBA30 and only 0.31 to 0.41 on MIRAGE novel families. The two numbers are consistent once the axis is stated.
The public benchmark harness accepts a predict(sequence, smiles) function or a predictions CSV and returns the family-support curve, G_m, matched-pK results, and bootstrap intervals. The code is on GitHub and the datasets are on Hugging Face.
One overall correlation hides the only regime that matters for a new target. MIRAGE reports both.