arXiv cs.LGPaper
Winning by Peeking: Unenforced Budgets and Test-Set Selection Inflate Short-Budget AutoML Comparisons
This is a useful methodological warning for anyone evaluating AutoML or benchmark claims generally: unenforced budgets and test-set peeking can manufacture a 78% win rate out of nothing. Treat vendor benchmark tables with the same skepticism this paper applies, especially any comparison run by the tool's own authors.