ROFT:只用自我复盘微调,Qwen3.5-4B 不靠强化学习也能提升智能体表现
Shockingly Simple Self-retrospection Improves Agentic Models Without RL
Qwen3.5-4B 只靠写复盘解释做微调,不用奖励模型就在 SWE-bench 上跑赢 GRPO,做智能体训练的可以看看这套 ROFT 流程。
论文提出 Retrospection-Only Fine-Tuning(ROFT):智能体完成任务后生成复盘解释,仅对解释文本做下一词预测训练,不需要外部教师或奖励信号。在 Qwen3.5-4B 上,ROFT 经 20 次更新后在 SWE-bench Verified 和 Pro 分别达到 49.2% 和 26.8% 的解决率,超过 GRPO 40 次更新后的 48.0% 和 25.3%。即使基础模型 64 次采样全部失败的任务,ROFT 也能学会解决。行为分析显示复盘会间接为动作分配信用,引导正确行为、抑制错误行为。
Shockingly Simple Self-retrospection Improves Agentic Models Without RL
People learn not only by repeating successful actions, but also by recounting and explaining their experiences, revising their understanding to guide future behavior. Can a language-model agent improve its future actions by training only on explanations of its own experience? We investigate this question by studying Retrospection-Only Fine-Tuning (ROFT), a minimal online procedure designed to isolate the effect of explanation-only training on subsequent behavior. The agent attempts a task, observes available feedback, generates a retrospective explanation, and is fine-tuned with a next-token prediction loss on the explanation tokens alone. The procedure uses neither an external teacher nor a reward-based policy update. In software-engineering experiments with Qwen3.5-4B, ROFT is trained on problems with mixed successful and unsuccessful base-model attempts. On held-out SWE-bench Verified and Pro, it reaches 49.2% and 26.8% solve rates after 20 updates without using a verifier, compared with GRPO's 48.0% and 25.3% after 40 updates in the evaluated runs, and makes faster early progress in training time and sampled attempts. It also learns to solve individual tasks on which all 64 sampled base-model attempts failed, showing that learning can begin without any initially successful trajectories. Behavioral analyses find that ROFT indirectly assigns credit to actions, encouraging good actions and discouraging incorrect ones. Moreover, prompting retrospections to emphasize more direct solutions yields shorter subsequent attempts even without an explicit length penalty. Together, these findings show that learning to explain can also improve learning to do, establishing self-generated retrospections as useful training targets and motivating further study of explanation-to-action transfer.