arXiv cs.CLPaper
Do Large Language Models Capture the Diversity in their Training Data?
Models are more confident and less diverse than their training distribution. This could explain why LLM outputs feel repetitive at scale, and why sampling strategies matter. It's a real observation but the implications for builders are unclear: do you want more diversity, or is confident output what you're actually paying for.