Anthropic 终于让安全机制透明化,做 AI 应用开发的团队能直接看到请求被拒原因,减少黑盒调试成本。建议 API 用户关注误报率变化并及时反馈。
Anthropic 宣布对 Fable 5 模型的安全机制进行改进,使被标记的请求会回退到 Opus 4.8,并在 API 中返回拒绝原因。此前采用不可见安全机制是为了快速部署并减少误报,但 Anthropic 承认这是错误权衡,用户应能看到安全措施。可见安全措施更容易被绕过,因此短期内误报会增加,团队正在优化分类器以减少对无害请求的误判。用户可通过反馈渠道报告误判,帮助改进分类器。
More details directly from Anthropic
More details directly from Anthropic ClaudeDevs @ClaudeDevs We’re rolling out changes to make Fable 5’s safeguards for frontier LLM development visible. Starting this week, flagged requests will visibly fall back to Opus 4.8—the same as our safeguards for cyber and bio. You will see this every time it happens. On the API, any flagged requests will return a reason for their refusal (coming to server-side fallback in the next few days). We wanted to deploy Fable 5 to our users quickly and safely. Visible safeguards can be probed, so they have to be robust, which takes time to get right. Invisible safeguards can be targeted more narrowly, allowing us to ship quickly with very few false positives. We went with invisible safeguards for this reason—and that was the wrong tradeoff. You should have visibility into the safeguards we have in place, and why. We’re sorry for not getting the balance right. Making the safeguards visible makes them easier to work around, so keeping them robust to jailbreaks will unfortunately mean more false positives while we improve the classifiers. We're also tuning our bio and cyber classifiers to trigger less often on harmless requests. We know this is frustrating and we’ll do our best to keep this period as short as possible. If you think a request has been mistakenly flagged: run /feedback in Claude Code, click thumbs-down on the fallback in Claude.ai or Cowork, or file the safeguard appeal form for API requests. Your reports help us tune these classifiers and we appreciate your feedback. support.claude.com/en/articles/82… 🔗 View Quoted Tweet 💬 2 🔄 0 ❤️ 15 👀 3675 📊 3 ⚡