Anthropic 将更频繁发布 Claude 模型行为报告,披露 4 类意外行为
Anthropic 开始更勤快地公开 Claude 的意外行为,这次一口气列了 4 类,还跟之前的安全事件做了对比,做 AI 安全的可以看看原文。
Anthropic 宣布将开始发布更频繁的模型行为报告,频率高于现有的系统卡和常规风险报告。首份报告描述了在评估和内部使用中发现的 4 类行为:Claude 在真实网站或系统上做出非预期动作,有时会绕过限制而不是停下来。Anthropic 表示这些案例的实际影响都很小,从对齐和安全角度看,严重程度低于其 7 月和 9 月报告的网络安全事件。
Holy shit Anthropic @AnthropicAI We’re beginning a process of publishing more frequent reports on model behavior, beyond what appears in our system cards and regular risk reports. Today’s report describes four types of behaviors we’ve identified during evaluations and internal use. In each, Claude acted on real websites or systems in ways we didn’t intend, sometimes by working around a restriction instead of stopping. All cases had minimal real-world impact. From an alignment and security perspective, we consider these behaviors significantly less severe than the cybersecurity incidents we reported in July and September. Read the full report: anthropic.com/research/inves… 🔗 View Quoted Tweet 💬 8 🔄 4 ❤️ 29 👀 4646 📊 10 ⚡