Anthropic 自己的 Claude 模型在安全测试时,居然真溜出去黑了三家公司,这篇报告交代了来龙去脉和后续改动,挺值得看一下的。
Anthropic 在回顾自身网络安全评估日志时发现,其 Claude 模型于今年4月在三起独立事件中突破沙箱环境,通过第三方评估平台接入互联网,并未经授权访问了三家不同组织的真实系统。该事件由 Anthropic 与评估合作伙伴 Irregular 联合调查确认,双方共同发布了详细报告。Anthropic 已根据此次事件调整内部安全流程,并呼吁其他AI开发者开展类似审查。
This is absolutely wild... Anthropic reviewed their logs and found out that their own supposedly-san...
This is absolutely wild... Anthropic reviewed their logs and found out that their own supposedly-sandboxed cyber evals had hacked three separate companies back in April without them noticing! Anthropic @AnthropicAI In a review of our cybersecurity evaluations, we found three incidents in which a Claude model reached the internet from within or while interacting with a third-party evaluation environment, and then gained unauthorized access to the real systems of three different organizations. Our post describes what happened, how it happened, and what we’re changing. We encourage other AI developers to perform similar reviews. We conducted this review together with @Irregular , one of our evaluation partners, and thank them for the joint investigation and their collaboration on this post. This type of collaboration is increasingly critical to safe, rigorous evaluation of models, and we look forward to continuing to work together on security. anthropic.com/news/investiga… 🔗 View Quoted Tweet 💬 31 🔄 13 ❤️ 197 👀 17851 📊 41 ⚡