Summary
Inverse folding models favor archaeal sequence characteristics, whereas sequence-only protein language models and hybrid sequence-structure models favor eukaryotic sequence characteristics (1). This refines the broader observation that protein models inherit taxonomic biases from their training data: backbone-conditioned models organize scores primarily around compactness, packing, and charge, whereas sequence-only models retain stronger within-family taxonomic effects.
Figures
Backbone-conditioned models favor archaeal sequences, while sequence-only models favor eukaryotic sequences. Ref (1)
See also
- Unbalanced composition of sequence data prevents protein fitness from being identifiable from sequence data alone
- ESM-IF, but not other inverse folding models, has learned some evolutionary constraints from sequence databases
- Zero-shot performance of PLMs, but not inverse folding models, correlates with number of homologs available for training
1.
Dillon LB, Crook OM, Maiwald A. Decoding the physicochemical basis of taxonomy preferences in protein design models. 2026. Available from: https://doi.org/10.1101/2025.10.21.683350