Summary
Protein language models cannot extrapolate to functional novelty or ultra-high-fitness variants. Zero-shot PLMs are mostly useful as coarse filters separating poor or unfit variants from fit variants, but do not reliably rank highly fit variants or prioritize new-to-nature functional novelty (1).
Figures


Figures from (1)
See also
- No one-size-fits-all best approach to zero-shot or few-shot protein fitness prediction
- Protein language models are better zero-shot predictors for ranking closely related sequences than distantly related sequences
- Fitness prediction
1.
Woolley PR, Feller AL, Ellington AD, Wilke CO. Overestimating zero-shot fitness prediction: Broad benchmarks mask local failures and practical limitations. openRxiv; 2026. Available from: https://doi.org/10.64898/2026.06.04.730121