Summary

PLM-based searches can recover remote homologs missed by conventional profile-HMM pipelines, particularly in the 20-25% sequence-identity regime (1,2). This does not mean that HMMs cannot identify remote homologs: HMM-HMM comparison in HHsearch was developed specifically for this task and detects relationships below 20% identity (3). (4) found that PLMs were better than HMM baselines at homolog detection at all scales (link).

1.
Kilinc M, Jia K, Jernigan RL. Improved global protein homolog detection with major gains in function identification. Proceedings of the National Academy of Sciences. 2023;120(9). Available from: https://doi.org/10.1073/pnas.2211823120
2.
Liu W, Wang Z, You R, Xie C, Wei H, Xiong Y, et al. PLMSearch: Protein language model powers accurate and fast sequence search for remote homology. Nature Communications. 2024;15(1). Available from: https://doi.org/10.1038/s41467-024-46808-5
3.
Söding J. Protein homology detection by HMM-HMM comparison. Bioinformatics. 2005;21(7):951–60. Available from: https://doi.org/10.1093/bioinformatics/bti125
4.
Wu KE, Chang H, Zou J. ProteinCLIP: enhancing protein language models with natural language. openRxiv; 2024. Available from: https://doi.org/10.1101/2024.05.14.594226