Summary

Aggregate benchmark correlations can mask weak within-category performance. Correlations computed across broad benchmark aggregates can exceed correlations within a biologically or experimentally meaningful subset, so global rank-correlation metrics can overstate practical local usefulness (1).

Figures

Ref (1)

1.
Woolley PR, Feller AL, Ellington AD, Wilke CO. Overestimating zero-shot fitness prediction: Broad benchmarks mask local failures and practical limitations. openRxiv; 2026. Available from: https://doi.org/10.64898/2026.06.04.730121