Stanford 与 Together AI 论文:给智能体分工互检优于辩论投票
Stanford 和 Together AI 的新论文,智能体别再投票吵架,学15道题分工互检,3个模型从48.8%干到66.7%。
Stanford 和 Together AI 的论文指出,让多个智能体辩论再投票,多数情况只是从已有答案里挑一个。他们让 1 个模型复盘团队历史对话,改写协作分工规则,比如谁负责复核、谁负责唱反调,仅用 15 道练习题就学会了这套玩法。在数学和物理测试上,3 个模型组成的团队平均得分 66.7%,而其中最强单模型只有 48.8%,还超过了在模型各自答案中完美挑选的上限,说明协作产生了任何单模型都没有的新正确答案。
Stop making your AI agents debate and vote.
New Stanford+Together AI paper shows teams that learn to check each other's work solve problems none of them got right alone.
Many agent setups just have models debate, then vote. That mostly picks an answer someone already had.
Here, 1 model reviewed past team chats and rewrote the team's playbook, like who double-checks whom and who plays devil's advocate. Learning this took just 15 practice problems.
On math and physics tests, the 3-model team averaged 66.7%, versus 48.8% for its best model alone. It even beat perfectly choosing among the models' own answers, so teamwork created right answers none of them had.
So skip the debate-and-vote script: give agents clear jobs like checker and challenger, and let past runs improve them.