SpaceXAI 发布 Grok 4.5,在代理知识工作基准上排名第四

Good model :)

精选理由

Grok 4.5 在代理知识工作基准上排第四,每个任务只要 49 美分,比前面的模型便宜 90%,性价比很能打。

AI 摘要

SpaceXAI 发布了 Grok 4.5,在 GDPval-AA v2 基准上以 Elo 1543 排名第四,仅次于 Anthropic 的最新 Claude 版本。每个任务的成本为 0.49 美元,低于 GLM-5.2 和 Kimi K2.6,且比排名更高的模型便宜近 90%。该模型在性能与成本的帕累托前沿上表现突出。

原文 · Sualeh Asif

Good model :)

Good model :) Artificial Analysis @ArtificialAnlys SpaceXAI just released Grok 4.5, and it ranks #4 on GDPval-AA v2 with an Elo of 1543 - behind only the latest Claude releases from Anthropic on real-world agentic knowledge work tasks Grok 4.5 achieved this score at a cost of $0.49 per GDPval task to sit clearly on the Pareto frontier for performance versus cost. This cost is lower than GLM-5.2 and Kimi K2.6, and nearly 90% cheaper than the models ahead of it on our leaderboard. We’re finalizing the remaining Artificial Analysis Intelligence Index evaluations and will share final results soon. Thanks to @SpaceXAI and @elonmusk for their collaboration testing this model ahead of release, and congratulations on the launch! 🔗 View Quoted Tweet 💬 0 🔄 0 ❤️ 42 👀 1492 📊 4 ⚡