Summary

Protein language models encode increasingly complex biological features as representations pass through deeper layers (1). Across ESM2 and AMPLIFY models, basic physicochemical properties and linear motifs are best captured in early layers, secondary structure in subsequent layers, and domain-level semantics in middle layers. Using sparse autoencoders, (2) likewise found that structure-level features, including disorder, emerge in later layers.

Figures

Ref (2)

See also

1.
Whitfield ST, Marty T, Vernon RM, Langmead CJ, Sridhar D, Fournier Q. High-resolution dissection of concept acquisition in different families of protein language models. 2026. Available from: https://doi.org/10.64898/2026.07.20.739599
2.
Adams E, Bai L, Lee M, Yu Y, AlQuraishi M. From Mechanistic Interpretability to Mechanistic Biology: Training, Evaluating, and Interpreting Sparse Autoencoders on Protein Language Models. openRxiv; 2025. Available from: https://doi.org/10.1101/2025.02.06.636901