Summary

Naively concatenating sequence and structural embeddings can underperform learned fusion mechanisms or sequence embeddings alone. Concatenating ESM-1b and GEARNet features was less effective than cross-attention, ESM-to-GEARNet, or ESM-1b alone (1). (2) similarly found that concatenation was outperformed by a ResNet autoencoder bottleneck. This does not imply that the modalities are inherently redundant: SPDesign obtained synergistic inverse-folding gains by combining structural sequence profiles, pretrained structural knowledge, and geometric features (3).

Figures

MethodF_maxAUPR
ESM-1b0.8640.889
ESM-GearNet
- w/ parallel fusion0.7330.759
- w/ series fusion0.8830.890
- w/ cross fusion0.8800.893

Ref (1)

See also

1.
Zhang Z, Wang C, Xu M, Chenthamarakshan V, Lozano A, Das P, et al. A Systematic Study of Joint Representation Learning on Protein Sequences and Structures. 2023; Available from: https://arxiv.org/abs/2303.06275
2.
Detlefsen NS, Hauberg S, Boomsma W. Learning meaningful representations of protein sequences. Nature Communications. 2022;13(1). Available from: https://doi.org/10.1038/s41467-022-29443-w
3.
Wang H, Liu D, Zhao K, Wang Y, Zhang G. SPDesign: protein sequence designer based on structural sequence profile using ultrafast shape recognition. Briefings in Bioinformatics. 2024;25(3). Available from: https://doi.org/10.1093/bib/bbae146