Summary
Measurements of protein-ligand effects (such as and ) carried out in different assays only weakly correlate with each other. The correlation is strengthened when the assays share many protein-ligand pairs (1). This means that substantial noise is introduced when combining data from multiple assays, even on the same protein. Another recent paper found that the same data collected in two different labs with different equipment has an average Spearman correlation of 0.73(2). Measurements of the same variants of the same protein from different labs can still correlate strongly, around Spearman 0.8 in examples from (3). Due to these factors, training deep-learning affinity predictors on many datasets from different sources therefore requires complex schemes (cite Boltz-2 paper)
Figures
Ref (1)
Ref(2)
Ref (3)
See also
- About 100k datapoints required to train accurate ddG predictor
- Spearman correlations of protein property prediction methods do not correlate perfectly with absolute error
- Kd differs from IC50, LD50, and GI50
- All-atom structure and affinity predictors partially generalize to point mutants