精选理由
ARC-AGI这个测试挺刁钻,GPT-5.6 Sol在没法记住之前推理的情况下学不会新游戏,说明长上下文对推理很重要。
ARC-AGI-3基准测试要求模型无需指令学习不熟悉的2D游戏。标准测试丢弃了GPT-5.6 Sol每步后的推理,并在上下文填满时丢弃早期动作。这导致模型无法积累有效经验,只能不断重启。测试结果凸显了长上下文维持对推理模型的关键性。
原文 · OpenAI
ARC-AGI-3 tests how well models can learn unfamiliar 2D games without instructions. The standard ha...
ARC-AGI-3 tests how well models can learn unfamiliar 2D games without instructions. The standard harness discarded GPT-5.6 Sol’s reasoning after each move and dropped earlier actions as the context filled up. The model had to keep starting over. Your browser does not support the video tag. 🔗 View on Twitter 💬 12 🔄 31 ❤️ 855 👀 91389 📊 87 ⚡