arXiv cs.CLPaper
Counting Documents Is Not Counting Text: Unit Bias in Web-PDF Corpus Statistics
This is a data quality catastrophe hiding in plain sight. If you've trained or fine-tuned on Common Crawl PDFs, your dataset is systematically biased toward short documents and missing more than half the available text in long ones. The TeX toolchain overrepresentation matters too. Go audit what you actually got versus what you thought you got.