Claude Opus 5 在部分基准上最大思考模式反而降低性能

An interesting thing I'm observing from the blog/system card is that on a good chunk of the reported...

精选理由

有人发现 Opus 5 开最大思考反而在某些基准上变差了,和直觉相反,值得研究。

AI 摘要

据观察,Claude Opus 5 的系统卡显示,在大约20-30%的已报告基准测试中,采用最大思考模式相比xhigh模式导致性能下降。通常增加思考时间和测试时计算会提升性能,但这些结果推翻了这一假设。这可能是小模型的自然涌现属性或后训练环节的问题。

原文 · Jerry Liu

An interesting thing I'm observing from the blog/system card is that on a good chunk of the reported...

An interesting thing I'm observing from the blog/system card is that on a good chunk of the reported benchmarks (~20-30% from a skim), Opus 5 max thinking leads to a degradation in performance compared to xhigh. Usually you would assume that as you increase thinking and test-time compute, performance goes up. Some of these results contradict that assumption. I wonder if this is a natural emergent property of smaller models or a posttraining issue. Claude @claudeai Introducing Claude Opus 5. It's a thoughtful and proactive model that comes close to the frontier intelligence of Fable 5 at half the price. Your browser does not support the video tag. 🔗 View on Twitter 🔗 View Quoted Tweet 💬 3 🔄 1 ❤️ 10 👀 1462 📊 4 ⚡

Claude Opus 5 在部分基准上最大思考模式反而降低性能 · AI 热点