Summary

MSA protein language models can identify intermolecular contacts and generalize from single-chain training to heterooligomeric protein-protein interactions. MSA Transformer extracts intermolecular contacts from correctly paired, but not incorrectly paired, MSAs (1). MSA Pairformer predicts protein-protein interface contacts and distinguishes binding from non-binding sequences despite being trained exclusively on individual chains (2). This extends the observation that structure prediction and design models trained on monomers can generalize to oligomers to MSA-based language models.

Details

The findings in MSA Pairformer were obtained by using a specific pairing strategy(2):

We integrated a proximity-based pairing scheme that infers genomic proximity directly from protein accessions. First, we translate UniProt and UniParc accessions into structured integers, based on the convention that sequentially numbered accessions reflect neighboring genes. Beginning from the highest scoring protein identified through an MMseqs2 search against the UniRef100, we iteratively select the closest unmatched protein within a predefined numerical threshold (default distance ≤ 20). This process is repeated until all suitable protein pairs have been identified. We allow either greedily pairing all possible matches or enforcing that all protein chains must be covered by paired database proteins.

Figures

MSA Pairformer scores distinguish high-fitness from low-fitness toxin-antitoxin interface sequences. Ref (2)

See also

1.
Lupo U, Sgarbossa D, Bitbol A-F. Pairing interacting protein sequences using masked language modeling. Proceedings of the National Academy of Sciences. 2024;121(27). Available from: https://doi.org/10.1073/pnas.2311887121
2.
Akiyama Y, Zhang Z, Tang O, Kim RS, Mirdita M, Steinegger M, et al. Expanding the scope of protein language modeling to protein-protein interactions with MSA Pairformer. Cell. 2026; Available from: https://doi.org/10.1016/j.cell.2026.06.029