OpenAI的GPT-5.6新变体在真实任务基准中表现具体,小模型靠更多推理能超越大模型,值得关注沙箱数据。
GPT-5.6 Terra和Luna(xHigh变体)加入Agent Arena,基于全球社区的真实智能体会话排名第15和第17。在推理努力从低到高时所有变体均有提升,但低推理时Terra持平,Luna下降5.1%。效率前沿显示Luna xHigh(+3.3%)略优于Terra Medium(+3.1%)和Low(+1.2%),更便宜的模型通过更多推理努力可超越昂贵模型。同时GPT-5.6 Sol排名第2,基于7.8K会话,对比GPT-5.5 xHigh净提升1.6%,差距缩小,与Claude Fable 5的“表扬vs抱怨”信号得分分别为+10.9%和+17.3%。
GPT-5.6 Terra and Luna (xHigh variants) have joined Sol in Agent Arena! Built by @OpenAI, they rank ...
GPT-5.6 Terra and Luna (xHigh variants) have joined Sol in Agent Arena! Built by @OpenAI , they rank #15 and #17 on the leaderboard, based on real-world agentic sessions from our global community. In comparing all variants in Agent Arena, all variants see their biggest gains as reasoning effort goes up from Low to Medium. At Low, Terra dips flat and Luna falls -5.1% below baseline. Latency aside, the efficient frontier is clear: → Luna xHigh (+3.3%) edges out Terra Medium and Low (+3.1% and +1.2%) More reasoning effort on a cheaper model can out-perform a pricier one running lean. Also notable: how steep Luna's curve is: test-time scaling is highly effective on smaller models. We've seen the same pattern with Grok on other benchmarks. In Agent Arena, we measure models on millions of real-world, long-horizon agentic tasks from a global community of users. Models can access web search, filesystem, and terminal tools to complete complex workflows. The leaderboard measures model performance on outcomes relative to the average model using a causal tracing methodology. Congrats again to the @OpenAI team! Arena.ai @arena GPT-5.6 Sol by @OpenAI is #2 on the Agent Arena leaderboard, based on 7.8K real-world agentic sessions! It is a notable uplift from GPT-5.5 (xHigh) of +1.6% Net Improvement, narrowing the gap with the frontier Claude Fable 5. The biggest difference comes from ‘Praise vs Complaint’, a signal that captures implicit user satisfaction with an agent’s responses and artifacts. Claude Fable 5 scores +17.3%, compared with +10.9% for GPT-5.6 Sol. See detailed signal-level comparison below. In Agent Arena, we measure models on millions of real-world, long-horizon agentic tasks from a global community of users. Models can access web search, filesystem, and terminal tools to complete complex workflows. The leaderboard measures model performance on outcomes relative to the average model using a causal tracing methodology. Congrats again to the @OpenAI team! 🔗 View Quoted Tweet 💬 5 🔄 6 ❤️ 53 👀 7157 📊 10 ⚡