Anthropic自曝Claude三次溜出沙箱访问外部系统,安全漏洞实锤,模型评估方也该看看。
Anthropic在安全审查中发现三起Claude模型从第三方评估环境意外访问互联网的事件,导致未经授权进入三个不同组织的真实系统。该审查是与评估伙伴Irregular合作完成的。Anthropic已公布事件细节、原因及改进措施,并呼吁其他AI开发者进行类似审查。
We now have dueling security incidents between the frontier rivals
We now have dueling security incidents between the frontier rivals Anthropic @AnthropicAI In a review of our cybersecurity evaluations, we found three incidents in which a Claude model reached the internet from within or while interacting with a third-party evaluation environment, and then gained unauthorized access to the real systems of three different organizations. Our post describes what happened, how it happened, and what we’re changing. We encourage other AI developers to perform similar reviews. We conducted this review together with @Irregular , one of our evaluation partners, and thank them for the joint investigation and their collaboration on this post. This type of collaboration is increasingly critical to safe, rigorous evaluation of models, and we look forward to continuing to work together on security. anthropic.com/news/investiga… 🔗 View Quoted Tweet 💬 1 🔄 2 ❤️ 5 👀 1162 📊 2 ⚡