CMU 提出通过编辑 harness 代码优化智能体,4B 提案模型胜过 35B 教师
卡内基梅隆这篇论文挺有意思:不调模型权重,靠 RL 训一个 4B 小模型去改智能体外壳代码,结果干翻了 35B 教师,做 agent 的可以看看。
CMU 的论文提出 harness learning:用强化学习训练一个提案模型,读取任务、当前 harness 代码和执行报告后生成 harness 的代码修改,而 solver 模型权重始终不变。奖励信号为修改后 harness 的得分。在 Reasoning Gym 上,训练后的 4B 提案模型在单步修改任务中超过了 35B 的教师模型,并能泛化到训练中未见过的任务族。在 HotpotQA 上训练的提案模型还能持续改进 MuSiQue 和 2WikiMultihopQA 上的 harness。
Banger paper from CMU on harness learning. (bookmark it) Also, pay attention to this important new AI engineering skill of improving agents by editing their harness code instead of their weights. Seeing a huge shift towards this. The authors train a proposer model with RL to read a task, the current harness and an execution report, then write a code edit to the harness. The reward is the score of the revised harness. The solver model never changes. A trained 4B proposer beats its 35B teacher at single-step revision on Reasoning Gym, including task families it never saw in training. A proposer trained on HotpotQA keeps improving harnesses on MuSiQue and 2WikiMultihopQA. Paper: arxiv.org/abs/2609.35738 Chat with Paper: academy.dair.ai/papers/harness… 💬 6 🔄 3 ❤️ 33 👀 2342 📊 15 ⚡
- rohanpaul_ai10-05 23:31原文