BehaviorTrace:检验 RL 训练数据归因方法的开放评测框架
Which Rollout Taught It That? BehaviorTrace and the Limits of Training-Data Attribution in Online RL
训练完 RL 模型想追查是哪条数据教会它某个行为?这篇直接证明现有归因方法大多被梯度大小和流畅度这类干扰项骗了,还给了套评测清单。
这篇论文研究在线 RL 微调后,能否找出教会语言模型新行为的具体训练 rollout。团队基于 GRPO 设计了植入行为的对照实验,发布开源评测框架 BehaviorTrace,在 Qwen2.5-1.5B 上跨三个种子测试。结果显示,仅按梯度大小排序的对照方法就达到随机基线的 4.2 到 4.5 倍,在三个种子中的两个上追平或超过最佳靶向估计器;在饱和检查点上,模型流畅度对行为标签的预测力不低于所有梯度方法。唯一在三个种子上都成立的信号是触发词梯度与行为发生位置构建的目标的对齐。论文还测试了 GAS(renormalized TracInCP)和 TRAK 风格估计器,并给出 RL 归因评测清单。
Which Rollout Taught It That? BehaviorTrace and the Limits of Training-Data Attribution in Online RL
When reinforcement learning teaches a language model a new behavior, can we find the training rollouts that taught it? And when an attribution method says it can, how do we know the answer is real? We study both questions on online RL fine-tuning with GRPO, using a planted behavior with a known cause. We release BehaviorTrace, an open evaluation harness that combines full-gradient sketching, the planted-behavior setup, and controls for gradient magnitude, fluency, headroom, and variation across seeds and generation draws. Across three seeds on Qwen2.5-1.5B, much of the apparent attribution signal comes from confounds. A control that ranks training steps by gradient size alone, with no behavior target, reaches 4.2 to 4.5 times chance and matches or beats the best targeted estimator on two of three seeds. At saturated checkpoints, model fluency predicts the behavior label at least as well as every gradient method we compared it with. Once fluency is controlled, the per-rollout results change from seed to seed and from one generation draw to the next, so a single run cannot settle the question. One signal does hold on all three seeds. The gradient of the trigger tokens aligns with a target built where the behavior actually occurs. We turn these findings into a checklist for evaluating attribution in RL. We test existing estimators, including GAS (renormalized TracInCP) and a TRAK-style estimator, and do not propose a new one.