Opus 5在ARC-AGI-3评估中得分是次优模型三倍

On ARC-AGI-3, an evaluation where AI models must solve novel problems, Opus 5’s score is three times...

精选理由

Opus 5在ARC-AGI-3上得分碾压其他模型,证明它的泛化和推理能力很强。

AI 摘要

在ARC-AGI-3基准上,Opus 5的得分是次优模型的3倍。该基准测试AI模型解决全新问题的能力,共包含400道题。Opus 5在多项类人推理指标上表现突出。

原文 · Claude

On ARC-AGI-3, an evaluation where AI models must solve novel problems, Opus 5’s score is three times...

On ARC-AGI-3, an evaluation where AI models must solve novel problems, Opus 5’s score is three times as high as the next best model. 💬 58 🔄 153 ❤️ 2852 👀 602482 📊 345 ⚡

Opus 5在ARC-AGI-3评估中得分是次优模型三倍 · AI 热点