Summary
Encoding germline and non-germline amino acids as separate token types improves antibody-language-model prediction of non-germline residues in CDRH3 without sacrificing strong framework prediction. PRISM uses a factorized 53-token vocabulary and separates germline from mutated residues in representation space (1).
See also
1.
Kim J, Blalock N, Kulkarni A, Nakamura K, Romero PA. Explicit representation of germline and non-germline residues improves antibody language modeling. openRxiv; 2026. Available from: https://doi.org/10.64898/2026.05.06.723387