Anthropic 的 Claude Opus 5 在真实编程和文本推理任务上跑赢了多数竞品,想看看新模型实力的可以关注这个榜单。
Arena AI 数据显示,Anthropic 的 Claude Opus 5 在启用 Max reasoning 时,在 Frontend Code Arena 中排名第 1,在 Text Arena(factuality on)中也排第 1。Opus 5 默认高推理模式在 Frontend Code Arena 中排第 3(仅次于 Kimi K3),在 Text Arena 中排第 2。这些结果覆盖 agentic 网页编程、文档推理和通用对话能力。其中 Opus 5 Max 的分数为初步数据,最终排名可能变化。
Arena AI说,这是Anthropic最新模型在真实任务上的数据:agentic网页编程、文档推理、通用对话能力都覆盖了。 不过Opus 5 Max的分数还是初步的,后续排名可能会有变动。 ht...
Arena AI说,这是Anthropic最新模型在真实任务上的数据:agentic网页编程、文档推理、通用对话能力都覆盖了。 不过Opus 5 Max的分数还是初步的,后续排名可能会有变动。 x.com/arena/status/2… Arena.ai @arena Exciting news: Claude Opus 5 with Max reasoning is #1 in the Frontend Code Arena and Text Arena with factuality on! Claude Opus 5 with default reasoning high is also very strong landing #3 in Frontend Code Arena, right behind Kimi K3 - and #2 in Text Arena (factuality on). This is real world data that @AnthropicAI 's newest model holds up on real world tasks: agentic web coding, document reasoning, and general chat capability. Claude Opus 5 Max’s score is still preliminary. We’ll continue to see how scores converge and share updates. Congrats to @AnthropicAI on the SOTA release! 🔗 View Quoted Tweet 💬 0 🔄 0 ❤️ 0 👀 188 ⚡