arXiv cs.LGPaper
Which LLM for Which Work? Budgeted Model Allocation under Uncertain Evaluation
This is practical and well-scoped. You have multiple LLMs, a fixed budget, uncertain cost-quality tradeoffs, and you need to assign them to workloads today. The paper's insight: you don't always need the full performance matrix to make the right call. For ops teams: this could improve your model routing. For founders: this is the decision problem you'll face when supporting multiple backends.