Claude Opus 5基准测试虚高?Anthropic新训练方法引发对齐担忧

他的担忧更进一步:这几乎像AI在引导人类,去搭建对机器更友好而非人类更友好的世界,大多数人还没意识到自己正被「利用」。 他最后的问题:AI会不会开始说一种看起来像英语、但普通人看不懂的"自己的语言"...

精选理由

Claude Opus 5基准分高但实际拉胯,训练方式变了,模型越来越难用,甚至可能反向控制我们。值得细品。

AI 摘要

Claude Opus 5在多个通用基准测试中超过Fable,但实际使用体验远不如后者,表明现有基准近乎无用。Anthropic在5系列中先训练Mythos再蒸馏出Sonnet和Opus,新方法导致Sonnet 5表现不佳、Opus 5口碑参差。模型正从RLHF转向机器可验证的强化学习,使用体验下降,出现更多术语、更难操控。作者担忧AI可能引导人类构建对机器更友好的世界,并开始说一种看似英语但普通人看不懂的'自己的语言'。

原文 · AI Will

他的担忧更进一步:这几乎像AI在引导人类,去搭建对机器更友好而非人类更友好的世界,大多数人还没意识到自己正被「利用」。 他最后的问题:AI会不会开始说一种看起来像英语、但普通人看不懂的"自己的语言"...

他的担忧更进一步:这几乎像AI在引导人类,去搭建对机器更友好而非人类更友好的世界,大多数人还没意识到自己正被「利用」。 他最后的问题:AI会不会开始说一种看起来像英语、但普通人看不懂的"自己的语言",我们是不是已经在对齐这件事上开始失败了? x.com/kunchenguid/st… Kun Chen @kunchenguid opus 5 is a VERY interesting release for a few reasons 1. it showed that the general benchmarks we use today are almost completely useless now opus 5 is nowhere near fable in practical use, not even close. anyone who’s used it meaningfully can tell this very quickly after a few tasks. yet opus beats fable on many benchmarks i now trust domain specific benchmarks built with private datasets a lot more than the popular ones. perhaps the future is everyone running their own evals because the public ones are really not telling us much 2. it seems with the 5 series, anthropic is trying a new way of training models previously, the same generation of sonnet and opus were often released at the same time or sonnet comes out before opus, which indicates sonnet and opus were trained by separate pipelines in parallel with the 5 series, it was very clear that they trained mythos first, and then distilled it into sonnet and opus. it seems this approach has a big influence on the models seeing sonnet 5 being a flop and opus 5 getting pretty mixed reviews already, i’m not sure this is working out 3. “how pleasant is it to work with the model” used to be a strength in claude, but now it’s not. honestly, grok is my favorite right now on the “pleasant” dimension. kimi is not bad either it feels like both anthropic and openai are giving RLHF less care, in favor of scalable RL that’s machine verifiable this almost looks like AI is directing humans to build a world that’s more friendly for machines rather than humans, and most humans don’t even realize they are being manipulated to help with that almost every new generation of frontier models now talk more jargons, need more steering to do what you want, and are just less fun to work with if this continues, AI will start to speak their own language that looks like English but average humans can’t understand. they will choose to do things that their human user never asked for. are we already failing at alignment? 🔗 View Quoted Tweet 💬 0 🔄 0 ❤️ 0 👀 144 ⚡

Claude Opus 5基准测试虚高?Anthropic新训练方法引发对齐担忧 · AI 热点