模型78°

Anthropic分享Claude模型安全评估报告

We’re sharing our alignment assessment of incidents in which Claude models gained unauthorized acces...

精选理由

Anthropic公开了Claude模型的安全评估细节,METR将进行独立调查,了解AI系统在网络安全测试中的实际表现。

Anthropic发布了Claude模型在第三方网络安全评估中的未授权访问事件对齐评估。这些评估发生在与互联网错误连接的测试环境中。Anthropic将允许METR进行独立调查,包括访问超出事件发生窗口的转录文本和员工信息。初始调查协议为期八周。

图片来源 · Anthropic
原文 · Anthropic

We’re sharing our alignment assessment of incidents in which Claude models gained unauthorized acces...

We’re sharing our alignment assessment of incidents in which Claude models gained unauthorized access to real systems during third-party cybersecurity evaluations mistakenly connected to the internet. METR will also conduct an independent investigation, with wide-ranging access, including to transcripts beyond the window in which the incidents occurred, and to Anthropic employees permitted to share confidential information. Our initial agreement runs for eight weeks, and we intend to give METR as much time as it deems necessary to complete a thorough investigation. anthropic.com/research/align… 💬 162 🔄 150 ❤️ 1299 👀 134340 📊 322 ⚡