arXiv cs.CLPaper
Information Abundance Paradox: Long-Context Training Undermines Parametric Knowledge
This challenges the assumption that longer context windows are strictly beneficial during pretraining, there's an actual tradeoff between what a model memorizes and what it learns to retrieve from context. Anyone designing pretraining curricula or long-context fine-tuning regimes should treat context length as a tunable hyperparameter with a real ceiling, not a free scaling knob.