SaProt is a protein language model trained on paired amino acid identities and Foldseek tokens derived largely from AlphaFold2 structures (1). It outperformed sequence-only baselines across ten downstream tasks, including zero-shot variant effect prediction.
Details
- Structural tokens for residues with pLDDT values less than 70 are masked.
See also
1.
Su J, Han C, Zhou Y, Shan J, Zhou X, Yuan F. SaProt: Protein Language Modeling with Structure-aware Vocabulary. In: International Conference on Learning Representations. 2024. Available from: https://openreview.net/forum?id=6MRm3G4NiU