Interpretability

What models internally encode and how they use it: attention analysis, layer-wise probing, sparse autoencoders, and mechanistic explanations. Generalization tests remain under evidence/generalization.

29 items with this tag.

Representation geometry

How similarities, clusters, distances, and manifolds in model representations relate to sequence, structure, or function. Practical feature extraction is a separate inference topic.

14 items with this tag.