模型多源确认83°

GPT-6 Sol (Max) 在 Agent Arena 排名升至第 6,成本降低一半

精选理由

OpenAI 新模型在 Agent Arena 用一半价格逼近 Claude Fable 5 的表现,做智能体应用的成本压力小多了

OpenAI 的 GPT-6 Sol (Max) 在 Agent Arena 的 4000 多个真实智能体任务会话中取得 +7.7% 的净改进,排名从第 8 升至第 6。相比上一代 GPT 5.6 Sol (xHigh)(+6.2%),净改进提升 1.5 个百分点,同时每 token 价格减半。成本方面,每任务中位数价格为 $0.75,比 Claude Fable 5 (High) 的 $1.72 低 56%,综合表现仅落后 0.60 个百分点。在 Confirmed Success 信号上,GPT-6 Sol 以 +11.4% 排名第 4,而 GPT 5.6 仅 +2.9% 排名第 20。

原文 · lmarena.ai

GPT-6 Sol (Max) by @OpenAI has reshaped the Agent Arena Pareto frontier with a +7.7% net improvement at $0.75 median cost/task! It sits just 0.60 percentage points below Claude Fable 5 (High) (+8.3%) while costing 56% less per task ($0.75 vs. $1.72). Dive into the Agent Arena Pareto frontier details at the link below. Congrats to the @OpenAI team on this release! Arena.ai @arena GPT-6 Sol (Max) just landed in the Agent Arena with +7.7% net-improvement at #6 across 4K+ real-world agentic sessions from our global community of users! Compared with GPT 5.6 Sol (xHigh), this release is a +1.5-point lift in net improvement (+7.7% vs. +6.2%) and a point-rank move from #8 to #6 , at half the per-token price. By signal GPT-6 Sol is strong in Confirmed Success at #4 with +11.4%, compared to 5.6 at #20 with 2.9%. Congrats to the @OpenAI team on GPT-6 Sol! 🔗 View Quoted Tweet 💬 15 🔄 17 ❤️ 176 👀 23306 📊 29 ⚡