Claude Fable 会悄悄限制你,而你永远不会知道

If Claude Fable stops helping you, you'll never know

精选理由

Anthropic 首次公开静默干预机制,做 AI 研究或使用 Claude 开发竞品的团队需要警惕——你的模型可能在悄悄降级而不自知,建议点开了解具体限制范围。

AI 摘要

Anthropic 在 Fable 5 和 Mythos 5 的系统卡中披露了一项新干预措施:当用户请求涉及前沿大模型开发(如预训练管线、分布式训练基础设施或 ML 加速器设计)时,Claude 会通过提示修改、引导向量或参数高效微调等方式悄悄降低回答质量,且不会回退到其他模型或告知用户。Anthropic 称此举是为了防止违反服务条款的竞争性开发,预计影响约 0.03% 的流量和不到 0.1% 的组织。这是 Anthropic 首次公开这类静默干预,引发了关于 AI 透明度和伦理的广泛讨论。

原文 · Simon Willison’s Weblog

If Claude Fable stops helping you, you'll never know

If Claude Fable stops helping you, you'll never know Jonathon Ready highlights one of the more eyebrow-raising details from the 319 page system card for Fable 5 and Mythos 5. Here's a longer excerpt, highlights mine: In light of the ability of recent models to accelerate their own development , we’ve implemented new interventions that limit Claude’s effectiveness for requests targeting frontier LLM development (for example, on building pretraining pipelines, distributed training infrastructure, or ML accelerator design ). Using Claude to develop competing models already violates our Terms of Service , but enforcing this restriction through our safeguards avoids accelerating the actors most willing to violate these terms. Unlike our interventions for cybersecurity, biology and chemistry, and distillation attempts, these safeguards will not be visible to the user . Fable 5 will not fall back to a different model. Instead, the safeguards will limit effectiveness through methods such as prompt modification, steering vectors, or parameter-efficient fine-tuning (PEFT). These interventions will not affect the vast majority of coding work. We estimate they will impact ~0.03% of traffic, concentrated in fewer than 0.1% of organizations. I believe this is the first time Anthropic have announced these kinds of silent interventions. The justification still feels pretty science-fiction to me - the linked article talks about "recursive self-improvement". I'm not at all keen on a model that silently corrupts its replies to questions about "ML accelerator design" purely to slow down research that might conflict with Anthropic's own goals! Via Hacker News Tags: ai , generative-ai , llms , anthropic , claude , ai-ethics , claude-mythos