OpenAI称GPT-5.6 Sol在定制ARC-AGI-3测试中超越Opus 5

OpenAI claims GPT-5.6 Sol beats Opus 5 on ARC-AGI-3 but only with its own custom test harness

精选理由

OpenAI推出GPT-5.6 Sol,在自定测试里赢了Anthropic的Opus 5,但换官方环境就掉到7.8%,差距挺大。

AI 摘要

OpenAI声称其GPT-5.6 Sol模型在ARC-AGI-3基准测试中通过自定义测试环境获得38.3%的分数,高于Anthropic Opus 5的30.2%。但在官方标准环境下,GPT-5.6 Sol仅得7.8%,远低于Opus 5的30.2%。OpenAI的定制测试保留了推理链和上下文压缩作为辅助手段,而Opus 5在无辅助下完成。这一差异引发了关于基准测试公平性的讨论。

原文 · Decoder

OpenAI claims GPT-5.6 Sol beats Opus 5 on ARC-AGI-3 but only with its own custom test harness

OpenAI counters Anthropic's ARC-AGI-3 record: GPT-5.6 Sol scores 38.3 percent, but only through its own API with retained reasoning and context compaction. In the official test environment, the model managed just 7.8 percent. Opus 5 hit its 30.2 percent without such aids. The article OpenAI claims GPT-5.6 Sol beats Opus 5 on ARC-AGI-3 but only with its own custom test harness appeared first on The Decoder .

OpenAI称GPT-5.6 Sol在定制ARC-AGI-3测试中超越Opus 5 · AI 热点