Summary
Fine-tuned protein language models generalize poorly to mutations at positions absent from the fine-tuning dataset (1). This is not universal: likelihood-based ranking fine-tuning outperformed embedding-based predictors on positional extrapolation splits, although performance still fell relative to random splits (2).
Figures

Ref (1); blue dots show position-based split
1.
Didi K, Alamdari S, Lu AX, Wittmann B, Johnston KE, Amini AP, et al. FLIP2: Expanding Protein Fitness Landscape Benchmarks for Real-World Machine Learning Applications. openRxiv; 2026. Available from: https://doi.org/10.64898/2026.02.23.707496
2.
Hawkins-Hooker A, Surana S, Simons J, Kmec J, Bent O, Duckworth P. Likelihood-based Fine-tuning of Protein Language Models for Few-shot Fitness Prediction and Design. openRxiv; 2024. Available from: https://doi.org/10.1101/2024.05.28.596156