Summary

Protein language models cannot extrapolate to functional novelty or ultra-high-fitness variants. Zero-shot PLMs are mostly useful as coarse filters separating poor or unfit variants from fit variants, but do not reliably rank highly fit variants or prioritize new-to-nature functional novelty (1).

Figures

Figures from (1)

See also

1.
Woolley PR, Feller AL, Ellington AD, Wilke CO. Overestimating zero-shot fitness prediction: Broad benchmarks mask local failures and practical limitations. openRxiv; 2026. Available from: https://doi.org/10.64898/2026.06.04.730121