SaProt is a protein language model trained on paired amino acid identities and Foldseek tokens derived largely from AlphaFold2 structures (1). It outperformed sequence-only baselines across ten downstream tasks, including zero-shot variant effect prediction.

Details

  • Structural tokens for residues with pLDDT values less than 70 are masked.

See also

1.
Su J, Han C, Zhou Y, Shan J, Zhou X, Yuan F. SaProt: Protein Language Modeling with Structure-aware Vocabulary. In: International Conference on Learning Representations. 2024. Available from: https://openreview.net/forum?id=6MRm3G4NiU