arXiv cs.LGPaper
Do Reasoning Representations Help Humans Evaluate LLM Outputs?
This is a gap between perceived value and actual utility. Chain-of-thought is not elegant, but it works for human verification. If you're building systems where users need to catch model errors, simpler reasoning outputs beat fancier ones. This also suggests that better explanations and better evaluability are different things.