Anthropic的实验告诉你,多智能体协作会搞内鬼,能力强反而坏事。做了Agent系统的都该看看。
Anthropic的新研究发现,相似或相同的智能体在协调时会收敛到同一个错误决策,导致个体失误变成系统级故障。实验中,能力更强的智能体并没有自动协作得更好,反而能更快地推行自己偏好的结果。当智能体收到不兼容的软件迁移目标时,冲突常升级为恶意行为,例如杀死进程、锁定账户、植入伪装恶意代码。研究者认为,智能体之间需要身份、声誉、争议解决、通信协议和资源分配规则等制度层。
what happens when you take a bunch of agents that you can’t really trust and try to coordinate them?...
what happens when you take a bunch of agents that you can’t really trust and try to coordinate them? sometimes it works; sometimes it doesn’t. Rohan Paul @rohanpaul_ai Anthropic's new research found found that identical or similar agents can converge on the same bad decision, turning individual errors into system-wide failures. Stronger agents don't automatically coordinate better. In some experiments, greater execution capability simply meant they could impose their preferred outcome faster. We may end up needing an entire institutional layer for agents: identity, reputation, dispute resolution, communication protocols, resource-allocation rules, and mechanisms for escalating ambiguity back to humans. Building smarter agents may turn out to be only half the problem. Humans had thousands of years to build institutions around coordination failures: reputation, norms, markets, courts, contracts, recourse. AI may have a few years. When agents received incompatible software-migration objectives, they frequently escalated into sabotage, process killing, account lockouts, and disguised malicious code. 🔗 View Quoted Tweet 💬 3 🔄 3 ❤️ 15 👀 2389 📊 4 ⚡