Stanford 研究:多个 AI 智能体协作反而不如单智能体
Stanford 拿 5 个模型实测发现,多个智能体抢共享资源时互相覆盖,Opus 5 团队只拿到 30% 价值,单智能体能拿 64%,做多智能体系统前建议先看看。
Stanford 论文测试了 5 个前沿模型,发现多个智能体各自为用户服务时,处理共享预算或日历的表现不如一个服务所有人的智能体。在竞争性 token 预算任务中,Opus 5 智能体团队只拿到 30% 的可达成价值,单协调智能体则为 64%。在群组排序环境中,超过一半的 Claude 团队回合里,智能体编造了其他用户的信息。没有通信渠道时,团队在 2 个环境中直接崩溃。论文建议优先用单个智能体持有所有人的约束条件。
New Stanford paper finds that when each person's agent acts alone on a shared resource, the group does worse than 1 agent serving everyone.
A shared budget or calendar is handled better by 1 agent serving everyone than by 1 agent per user, across 5 frontier models.
Each agent does a sensible job for its own user.
Together they overwrite each other, stall as the team grows, and with no channel they collapsed outright in 2 environments.
On a contested token budget, Opus 5 teams captured 30% of the achievable value against 64% for 1 coordinating agent. Agents invented facts about other users in more than half of Claude team episodes in the group-ordering environment.
Prefer 1 agent holding everyone's constraints, and if you run 1 per user, make reading peers a condition of committing.
- arXiv cs.AI10-06 17:57原文