Summary

Random data splits overestimate the ability of general-purpose protein language model embeddings to predict nanobody expression because the embeddings encode antibody-program identity rather than transferable determinants of expression (1). General-purpose ESM2 embeddings achieved a ROC-AUC of approximately 0.88 under a random split but only 0.68 when evaluated on held-out antibody programs. AINN-P1 and domain-fine-tuned ESM2 retained ROC-AUC values of 0.81 and 0.83, respectively, on unseen programs. This qualifies comparisons such as Antibody LMs are worse for expression prediction than generic PLMs: model rankings and apparent representation quality can depend strongly on whether the evaluation split separates related discovery programs.

Figures

Leave-program-out evaluation reveals substantially more leakage in general-purpose ESM2 representations than in AINN-P1 or domain-fine-tuned ESM2 representations. Ref (1)

See also

1.
Wang R, Jin K, Pan L. AINN-Express: A Leakage-Aware, Sequence-Only Predictor of VHH Antibody Expression Built on the AINN-P1 Protein Foundation Model. 2026. Available from: https://doi.org/10.64898/2026.07.21.739256