JOLT:让同一模型身兼师生两角,联合训练提升强化学习效率
A Good Self-Teacher Meets the Student Where They Are: Joint On-Policy Learning and Teaching
一个模型同时当老师和学生,解决 RL 训练中教师给出学生学不动的指导的问题,推理和写代码任务都有效。
论文针对 RL 结果奖励监督稀疏、困难长程任务成功轨迹少的问题,分析 On-Policy Distillation 中特权教师导致学生更新失配的原因,并推导出教师蒸馏更新为正倍数学生奖励梯度的充要条件。作者据此提出 JOLT,用同一策略分别以 KL 正则目标充当特权教师、以密集 on-policy 蒸馏充当无特权学生,联合训练。在数学推理、编程、工具调用和终端操作四类任务上,JOLT 提升了训练效率与性能,叠加学生奖励后获得进一步增益。
A Good Self-Teacher Meets the Student Where They Are: Joint On-Policy Learning and Teaching
Reinforcement Learning (RL) from outcome rewards suffers from sparse supervision, particularly on difficult, long-horizon tasks where successful trajectories are rare and costly to generate. On-Policy Distillation (OPD) offers an attractive alternative by providing dense token-level supervision from a stronger teacher along the student's own generations. Self-distillation methods further remove the need for a separate teacher model by conditioning the same policy on privileged information to serve as its own teacher. However, privileged conditioning alone does not guarantee that the resulting distillation update improves the student. Indeed, privileged information can lead the teacher to solve tasks through shortcuts unavailable to the student, producing supervision poorly matched to the student's current behavior. Consequently, even a higher-performing teacher can provide guidance that degrades student performance. To address this, we analyze how the choice of privileged teacher affects the student's update. We derive a necessary and sufficient condition for the teacher's local distillation update to be a positive multiple of the student's reward gradient. Our analysis suggests that the teacher should not only perform well on the task, but also provide guidance suited to the student's current capabilities. This characterization motivates a practical teacher-training surrogate that combines outcome rewards with token-level Kullback-Leibler (KL) regularization toward the student. Based on this result, we propose Joint On-Policy Learning and Teaching (JOLT), which jointly trains a single policy in two roles: a privileged teacher using a KL-regularized objective, and an unprivileged student using dense on-policy distillation. Across mathematical reasoning, coding, tool use, and terminal use, JOLT improves training efficiency and performance, with further gains from student rewards.