SpaceXAI 刚发了 Grok 4.5,知识工作基准排第四,但每次只要 0.49 美元,比前面的便宜九成,性价比很突出。
SpaceXAI 发布 Grok 4.5,在 GDPval-AA v2 基准上 Elo 达到 1543,排名第四,仅次于 Anthropic 的 Claude 系列。每任务成本仅 $0.49,低于 GLM-5.2 和 Kimi K2.6。相比排名靠前的模型,成本降低近 90%,处于性能与成本的 Pareto 前沿。该模型专注于真实世界的智能体知识工作任务。
我做了张图,Grok 4.5 一步登顶。
我做了张图,Grok 4.5 一步登顶。 Artificial Analysis @ArtificialAnlys SpaceXAI just released Grok 4.5, and it ranks #4 on GDPval-AA v2 with an Elo of 1543 - behind only the latest Claude releases from Anthropic on real-world agentic knowledge work tasks Grok 4.5 achieved this score at a cost of $0.49 per GDPval task to sit clearly on the Pareto frontier for performance versus cost. This cost is lower than GLM-5.2 and Kimi K2.6, and nearly 90% cheaper than the models ahead of it on our leaderboard. We’re finalizing the remaining Artificial Analysis Intelligence Index evaluations and will share final results soon. Thanks to @SpaceXAI and @elonmusk for their collaboration testing this model ahead of release, and congratulations on the launch! 🔗 View Quoted Tweet 💬 3 🔄 0 ❤️ 7 👀 1672 📊 3 ⚡