研究:RL 后训练哪些题会变好,旧 checkpoint 比先验信号预测更准
What Will Post-Training Fix? Per-Problem Gains Are Shared Across Independent RL Runs, and Existing Checkpoints Predict Them Better Than A Priori Signals
arXiv 上这篇挺有意思:想知道 RL 训练后哪些题能解出来,用别的模型训过的 checkpoint 预测,比一堆先验指标都准。
研究者在 DeepSeek-R1-Distill-Qwen-1.5B 和 Qwen2.5-Math-1.5B 两个基座上跑了 18 次后训练,在最多 1532 道竞赛数学题上逐题评估。独立 RL 运行对哪些低通过率题目会改善高度一致,噪声天花板约 0.9,但两个基座之间相关性只有 rho=0.25,说明这种共识属于基座模型而非题目本身。先验信号(基座通过率、正确解似然、更大模型通过率及其组合)只能解释 0.30 和 0.24 的可解释方差。而来自另一模型家族的单个 checkpoint 的逐题增益预测新运行达 0.52,明显优于组合先验信号的 0.33。
What Will Post-Training Fix? Per-Problem Gains Are Shared Across Independent RL Runs, and Existing Checkpoints Predict Them Better Than A Priori Signals
Data selection, curricula and the evaluation of post-training recipes all assume that we can tell, before training, which problems a model will improve on. We test this assumption directly. For two base models, DeepSeek-R1-Distill-Qwen-1.5B and Qwen2.5-Math-1.5B, we evaluate eighteen post-training runs on up to 1532 competition math problems with many samples per problem, and compare signals available before training against a noise ceiling derived from the agreement between disjoint subsets of runs. Three findings hold on both base models. What post-training fixes is shared: independent runs agree on which rarely solved problems improve, with a noise ceiling of about 0.9, yet the two base models agree with each other only at rho=0.25 - the shared component belongs to the base model, not to the problem. A priori signals capture a minority of it: base pass rate, the likelihood of a correct solution, a larger model's pass rate and their combinations explain only 0.30 and 0.24 of the explainable variance in gains. Existing checkpoints are the better predictor: the per-problem gains of a single checkpoint from another family predict a new run better than every a priori signal, alone or combined (0.52 vs. 0.33 and 0.34 vs. 0.17 against the combined signals). The conclusions hold on problems from 2025-2026 competitions and when the baselines on the two sides of every comparison are estimated independently. We propose the noise ceiling as a standard companion to per-problem signals.