arXiv cs.CLPaper
Test-Time Scaling in the Wild: Why Exploitation, Not Exploration, Is the Bottleneck
This matters if you've been throwing inference budget at reasoning models for non-verifiable tasks like legal or medical drafting and wondering why gains plateau. The fix isn't more sampling, it's better selection and reward modeling on the output side. Anyone building agents for fuzzy domains should read the decomposition before tuning TTS knobs further.