Summary
Confidence metrics from structural modeling and design neural networks are not calibrated, target-independent predictors of binding affinity for de novo binders (1,2). This includes pLDDT, ipTM, Rosetta scores, ProteinMPNN likelihoods, and ESM-2 log-likelihoods. They can nonetheless contain target-specific classification or ranking signal: PAE weakly correlates with antibody-antigen affinity and better separates binders from nonbinders (3), while AlphaFold3 ipTM distinguishes antibody binders on some targets (4,5). Binder classification or target-specific enrichment should therefore not be conflated with general affinity regression. The limitation of pLDDT is also broadly true of monomeric protein stability. Metrics like ipTM can quickly respond to even small changes in the conditioning representations used to drive diffusion-based structure prediction, showing how brittle such metrics are with respect to input sequences and MSAs (6).
Figures
Ref (1)
See also
- Protein folding neural networks cannot predict protein stability
- Protein structure prediction and design metrics don’t correlate with expression probability
- Sequence- and structure-derived ML quality metrics from ML do not correlate with each other