Summary
Naively concatenating sequence and structural embeddings can underperform learned fusion mechanisms or sequence embeddings alone. Concatenating ESM-1b and GEARNet features was less effective than cross-attention, ESM-to-GEARNet, or ESM-1b alone (1). (2) similarly found that concatenation was outperformed by a ResNet autoencoder bottleneck. This does not imply that the modalities are inherently redundant: SPDesign obtained synergistic inverse-folding gains by combining structural sequence profiles, pretrained structural knowledge, and geometric features (3).
Figures
| Method | F_max | AUPR |
|---|---|---|
| ESM-1b | 0.864 | 0.889 |
| ESM-GearNet | ||
| - w/ parallel fusion | 0.733 | 0.759 |
| - w/ series fusion | 0.883 | 0.890 |
| - w/ cross fusion | 0.880 | 0.893 |
Ref (1)
See also
1.
Zhang Z, Wang C, Xu M, Chenthamarakshan V, Lozano A, Das P, et al. A Systematic Study of Joint Representation Learning on Protein Sequences and Structures. 2023; Available from: https://arxiv.org/abs/2303.06275
2.
Detlefsen NS, Hauberg S, Boomsma W. Learning meaningful representations of protein sequences. Nature Communications. 2022;13(1). Available from: https://doi.org/10.1038/s41467-022-29443-w
3.
Wang H, Liu D, Zhao K, Wang Y, Zhang G. SPDesign: protein sequence designer based on structural sequence profile using ultrafast shape recognition. Briefings in Bioinformatics. 2024;25(3). Available from: https://doi.org/10.1093/bib/bbae146