Are You Learning Biological Signal or Shortcuts? Auditing and Mitigating Bias in Protein-Protein Interaction Datasets
This matters if you're building biotech ML, because it forces you to rethink how you split PPI data and what signals your model is actually learning. The paper audits three major PPI databases and identifies previously unreported topological shortcuts, which means your current training pipeline is probably contaminated. Fix your data curation before you trust your model's predictions on unseen proteins.