arXiv cs.LGPaper
Scaling Near-Optimal SFT-RL Annotation Budget Allocation from Small to Large LLMs
Stop searching for a single magic SFT-RL ratio. The paper shows you can find a wide band of good allocations by testing on a small proxy model, then apply it to your production model without retuning. This saves you from running expensive large-model ablations. Practical and immediately usable for anyone doing post-training.