Summary

Fine-tuning PLMs ESM2, ProtT5, and Ankh virtually always improves property prediction (variant effect prediction, stability prediction, function prediction, others) compared to zero-shot across eight supervised benchmarks (1). However, this is not a guarantee for every objective or evaluation split: fine-tuning can degrade representations or generalization through high variance and catastrophic forgetting (2,3). When fine-tuning on paired antibody sequences, retaining the unpaired pretraining data and objective mitigated forgetting (4). Outside protein ML, RLHF also worsened the performance of GPT-4 on some tasks (5).

Figures

Figure 1 from (1)

Ref (6)

Ref (2)

See also

1.
Schmirler R, Heinzinger M, Rost B. Fine-tuning protein language models boosts predictions across diverse tasks. Nature Communications. 2024;15(1). Available from: https://doi.org/10.1038/s41467-024-51844-2
2.
Detlefsen NS, Hauberg S, Boomsma W. Learning meaningful representations of protein sequences. Nature Communications. 2022;13(1). Available from: https://doi.org/10.1038/s41467-022-29443-w
3.
Heinzinger M, Weissenow K, Sanchez JG, Henkel A, Mirdita M, Steinegger M, et al. Bilingual language model for protein sequence and structure. NAR Genomics and Bioinformatics. 2024;6(4). Available from: https://doi.org/10.1093/nargab/lqae150
4.
Kenlay H, Dreyer FA, Kovaltsuk A, Miketa D, Pires D, Deane CM. Large scale paired antibody language models. PLOS Computational Biology. 2024;20(12):e1012646. Available from: https://doi.org/10.1371/journal.pcbi.1012646
5.
Chen L, Zaharia M, Zou J. How Is ChatGPT’s Behavior Changing Over Time? Harvard Data Science Review. 2024;6(2). Available from: https://doi.org/10.1162/99608f92.5317da47
6.
Jiang K, Yan Z, Bernardo MD, Sgrizzi SR, Villiger L, Kayabolen A, et al. Rapid protein evolution by few-shot learning with a protein language model. openRxiv; 2024. Available from: https://doi.org/10.1101/2024.07.17.604015