Summary
Protein language models can behave as in-context learners, but tend to retrieve evolutionary patterns from supplied homologs rather than reason causally about sequence–fitness relationships. This can distort the relationship between model likelihood and biological fitness (1).
See also
- Protein language models are better zero-shot predictors for ranking closely related sequences than distantly related sequences
- PLMs downweigh probability of sequences with multiple mutations
1.
Kantroo P, Wagner GP, Machta BB. In-Context Learning can distort the relationship between sequence likelihoods and biological fitness. 2025; Available from: https://arxiv.org/abs/2504.17068