GPT-6 Luna (Max)登上Agent Arena前沿
OpenAI发布GPT-6 Luna (Max),成本大幅降低,性能接近高端模型,性价比突出。
GPT-6 Luna (Max)在Agent Arena前沿实现了+1.59%的净改进,中位任务成本仅为0.05美元。该模型比DeepSeek-V4.1-Flash成本低29%,比Hy4 preview成本低76,净改进分数仅落后2.27和2.51个百分点。与GPT-6 Sol (Max)相比,GPT-6 Luna (Max)成本低94%,与GPT-6 Astra (Max)相比成本低98%。
Update to the Agent Arena Pareto frontier: GPT-6 Luna (Max) has landed on the Agent Arena Pareto frontier with +1.59% net improvement at a median cost of just $0.05 per task! Luna costs 29% less than DeepSeek-V4.1-Flash and 76% less than Hy4 preview, while trailing their net-improvement scores by just 2.27 and 2.51 percentage points respectively. Compared to other GPT models on the Pareto frontier, GPT-6 Luna (Max) costs 94% less than GPT-6 Sol (Max), and 98% less than GPT-6 Astra (Max). Congrats to the @OpenAI team on this cost-efficient release! Arena.ai @arena GPT-6 Luna (Max) by @OpenAI is #23 in Agent Arena with +1.6% net improvement across 8K real-world agentic sessions from our global community of users. Although GPT-6 Luna (Max) did not land on the Agent Arena Pareto frontier, it remains a cost-efficient model. Its $0.05 median cost per task is 94% lower than GPT-6 Sol (Max) at $0.82 and 98% lower than GPT-6 Astra (Max) at $2.59. Its net-improvement score also comes within 0.09 percentage points of #22 GPT 5.5, while costing 91% less than its $0.56 median cost per task. This release is a six-place point-rank move over GPT-5.6 Luna (xHigh), at -0.9% and #29 ! By signal, GPT-6 Luna’s clearest gains over GPT-5.6 Luna are in: - Confirmed Success: #17 (+4.6%) vs. #33 (-4.6%) - Bash Recovery: #21 (+4.2%) vs. #27 (+2.2%) Congrats to the @OpenAI team on this release! 🔗 View Quoted Tweet 💬 10 🔄 3 ❤️ 34 👀 7142 📊 10 ⚡