SpaceXAI发布Grok-4.6模型

Grok-4.6 (xHigh) by @SpaceXAI has landed in Agent Arena, with its strongest category result in Code:...

精选理由

SpaceXAI的Grok-4.6在代码任务上表现亮眼,确认成功率提升15%,成本与Claude Opus 4.6相当。

AI 摘要

Grok-4.6(xHigh)在Agent Arena代码类别中排名第12,净提升7.7%,基于4500+真实智能体代码会话。在确认成功率方面排名第6,提升15%。整体确认成功率提升13.2%,排名第6。与Grok-4.5相比有显著提升,中位任务成本为1.12美元。

原文 · lmarena.ai

Grok-4.6 (xHigh) by @SpaceXAI has landed in Agent Arena, with its strongest category result in Code:...

Grok-4.6 (xHigh) by @SpaceXAI has landed in Agent Arena, with its strongest category result in Code: ranking #12 with a +7.7% net improvement, based on 4.5K+ real-world agentic code sessions. Within Code, Grok-4.6 also landed #6 in Confirmed Success, with +15%! When looking at Agent Arena overall, its clearest strength by signal is Confirmed Success (an explicit “yes, that worked” from users), with +13.2% net improvement overall ( #6 ). This is a significant jump from Grok-4.5 with +5% ( #18 ). With a median cost per task of $1.12, it is on par with Claude Opus 4.6 at $1.19, but doesn’t quite make the Pareto Frontier based on its performance. Grok-4.6 ranks #15 overall, with +6.1% net improvement. See more info by signal overall below. It also landed #12 in Chat (+5.0%) and #17 in Work (+6.1%). Congrats to the @SpaceXAI team on this release! SpaceXAI @SpaceXAI Introducing Grok 4.6. It delivers frontier intelligence and is a significant improvement over Grok 4.5 at the same price. 🔗 View Quoted Tweet 💬 6 🔄 0 ❤️ 22 👀 4479 📊 6 ⚡

SpaceXAI发布Grok-4.6模型 · AI 热点