arXiv cs.CLPaper
What Aggregate Scores Miss: Measuring Item-Level Regressions in Commercial LLM API Migrations
Aggregate scores lie. You can be told a model is better overall while specific capabilities you depend on get worse. If you're migrating to a new API version, don't trust the headline numbers. Run your actual workload against both models at scale and measure item-level deltas. This is not academic: it's a production decision-making tool.