Summary

Sparse autoencoders trained on protein language model representations recover interpretable features associated with protein families, gene names, and Gene Ontology terms. These features can be extracted from both pooled protein representations and residue-level representations (1).

See also

1.
Gujral O, Bafna M, Alm E, Berger B. Sparse autoencoders uncover biologically interpretable features in protein language model representations. Proceedings of the National Academy of Sciences. 2025;122(34):e2506316122. Available from: https://doi.org/10.1073/pnas.2506316122