精选理由
Opus 5在ARC-AGI-3上得分碾压其他模型,证明它的泛化和推理能力很强。
在ARC-AGI-3基准上,Opus 5的得分是次优模型的3倍。该基准测试AI模型解决全新问题的能力,共包含400道题。Opus 5在多项类人推理指标上表现突出。
原文 · Claude
On ARC-AGI-3, an evaluation where AI models must solve novel problems, Opus 5’s score is three times...
On ARC-AGI-3, an evaluation where AI models must solve novel problems, Opus 5’s score is three times as high as the next best model. 💬 58 🔄 153 ❤️ 2852 👀 602482 📊 345 ⚡