Summary
ESM-IF, but not other inverse folding models, has learned some evolutionary constraints from sequence databases (1). This is evidenced by a dependence on the availability of training data that is comparable to that of protein language models and Potts models. A practical example is in RNA-guided nucleases, where ESM-IF but not other inverse folding models introduced catalytic residues necessary for function (2).
Figures

Ref (1)
See also
- Zero-shot performance of PLMs, but not inverse folding models, correlates with number of homologs available for training
- Inverse folding models favor archaeal proteins while sequence-only PLMs favor eukaryotic proteins
- Inverse folding selects more evolutionarily conserved design positions than sequence-only protein language models
1.
Li F-Z, Yang J, Johnston KE, Gürsoy E, Yue Y, Arnold FH. Evaluation of machine learning-assisted directed evolution across diverse combinatorial landscapes. Cell Systems. 2025;16(9):101387. Available from: https://doi.org/10.1016/j.cels.2025.101387
2.
Skopintsev P, Esain-Garcia I, DeTurk EC, Yoon PH, Zhou Z, Weiss T, et al. Structure and evolution-guided design of minimal RNA-guided nucleases. Science. 2026 Jul;393(6808):313–8. Available from: http://dx.doi.org/10.1126/science.aed6123