论文73°

ROFT提升智能体模型无需强化学习

精选理由

ROFT方法让智能体通过自我解释就能提升性能,比强化学习简单多了。

研究人员提出回顾式微调(ROFT)方法,通过仅对智能体自身经验解释进行训练来改进其后续行为。该方法让智能体尝试任务后观察反馈,生成回顾性解释,并仅对解释令牌进行下一令牌预测损失微调。实验表明这种简单方法能有效提升智能体性能,无需复杂强化学习。

原文 · Tanishq Abraham (论文推介)

Shockingly Simple Self-retrospection Improves Agentic Models Without RL

"Can a language-model agent improve its future actions by training only on explanations of its own experience? We investigate this question by studying Retrospection-Only Fine-Tuning (ROFT), a minimal online procedure designed to isolate the effect of explanation-only training on subsequent behavior. The agent attempts a task, observes available feedback, generates a retrospective explanation, and is fine-tuned with a next-token prediction loss on the explanation tokens alone."

link: https://t.co/x3PiRSc1XX