精选理由
这篇帖子讲了个后训练小诀窍:短任务训练就能让智能体扩展到8-32倍的长周期,省时省力,值得试试。
该帖子提出一个后训练技巧:训练长周期智能体(long-horizon agents)时无需使用长周期 rollout。可以在短任务(short tasks)上训练,这类任务奖励成本低且验证容易。然后利用框架(harness)将能力扩展8-32倍,显著降低训练成本。
原文 · elvis
Something really neat about this work that might not get picked up easily: for post-training, you ma...
Something really neat about this work that might not get picked up easily: for post-training, you may not need long-horizon rollouts to train long-horizon agents. You could train on short tasks where rewards are cheap and verification is easy, and let the harness carry it 8-32x further. 💬 0 🔄 0 ❤️ 1 👀 436 ⚡