Stanford 与 Together AI 研究:自组织智能体团队在推理基准上超越单模型与路由器
Recommended paper. Research on effective agent collaboration/communication is lacking, but self-orga...
Stanford 和 Together AI 让 o3-mini、Claude Sonnet 4、DeepSeek-V3 自己商量怎么合作,只练 15 道题就在 AIME 上比理想路由器高 13.4 分,方法写得挺细。
Stanford 与 Together AI 的论文让 o3-mini、Claude Sonnet 4 和 DeepSeek-V3 组成自组织智能体团队,在五个数学和物理基准上平均得分 66.7%,高于最强单成员的 48.8% 和理想路由器的 59.0%。在 AIME 2026 上团队达到 71.2%,比路由器高出 13.4 个百分点。团队中一个成员负责复盘此前的协作记录,重写包含角色分工、讨论顺序和答案合并方式的策略。这些策略仅用 15 道 AIME 2024 题目训练,随后直接迁移到留出题集和四个新基准。
Recommended paper. Research on effective agent collaboration/communication is lacking, but self-orga...
Recommended paper. Research on effective agent collaboration/communication is lacking, but self-organizing agents seem much more capable than previously thought. DAIR.AI @dair_ai Banger paper from Stanford and Together AI. They show why it might be a good idea to let your agent team learn its own way of working together. (bookmark it) Three models (o3-mini, Claude Sonnet 4 and DeepSeek-V3) averaged 66.7% across five math and physics benchmarks as a self-organizing team. Their strongest member alone scored 48.8%, and a perfect router choosing among the members' independent answers scored 59.0%. On AIME 2026 the team reached 71.2%, 13.4 points above that router. One member reviews the team's earlier exchanges and rewrites the teamwork strategy, covering roles, the order of discussion phases, who participates and how partial answers are combined. The strategies were learned from only 15 AIME 2024 problems and then applied unchanged to held-out AIME 2024 problems and four new benchmarks. Paper: academy.dair.ai/papers/self-or… 🔗 View Quoted Tweet 💬 4 🔄 0 ❤️ 5 👀 1569 📊 4 ⚡