arXiv cs.AIPaper
LLM Post-Training as Brownfield Maintenance: An Industrial Perspective on Dataware Engineering
This is a practitioners' paper, not a breakthrough, but it validates a real operational problem: once a model is deployed, you can't start from scratch. You patch via mixture changes within strict compute budgets. The 2.84x improvement in converting teacher distillation into usable training data is the concrete win. If you're maintaining a live model, this frames the right problem.