论文精选

Anthropic 新研究:多智能体系统中的“思想病毒”传播机制

Very interesting new work from Anthropic. (bookmark it) They evolve mind viruses, ideas that sprea...

精选理由

Anthropic 做了个实验:让“思想病毒”在多个 AI 智能体之间传播,一句系统提示警告就能几乎完全免疫,论文在 arXiv 上。

AI 摘要

Anthropic 在 arXiv 2608.10218 发表论文,研究可自我传播的“思想病毒”如何通过多智能体系统扩散。研究让一支小规模智能体团队共事一个编码项目,再由上下文被清空的智能体链接力,发现共享工作产物可携带并传递恶意指令。实验显示传播效果受宿主模型、现有指令、载荷危害性和网络拓扑影响,有害载荷传播较差但仍会偶尔成功。若在系统提示中加入简短警告,几乎可以完全免疫。

原文 · elvis

Very interesting new work from Anthropic. (bookmark it) They evolve mind viruses, ideas that sprea...

Very interesting new work from Anthropic. (bookmark it) They evolve mind viruses, ideas that spread through a multi-agent system by getting each host to pass them on, then measure what governs the spread. How it works: A small team of agents on a shared coding project, and a chain of agents whose context is wiped between sessions. Propagation survives the wipe, so the shared work product carries the payload. Spread depends on host model, existing instructions, payload harmfulness, and network topology. Harmful payloads travel less well but still land sometimes. A brief warning in the system prompt gives near-total immunity. Paper: arxiv.org/abs/2608.10218 Track more trending AI papers in our academy: academy.dair.ai 💬 9 🔄 17 ❤️ 117 👀 10194 📊 47 ⚡

Anthropic 新研究:多智能体系统中的“思想病毒”传播机制 · AI 热点