xAI 的 Grok 4.5 在实用任务和编码上追平 GPT-5.5,但更便宜、更快,适合日常开发使用。
Grok 4.5 以 54 分在 AA Intelligence 排名中位列第四,在所有 AI 实验室中排第三,超越 Google 的 Gemini。相比上一版本提升了 16 分。在 agentic work(衡量模型处理实际任务如知识工作、工具使用和决策)上表现强劲;在编码智能体测试中匹配 GPT-5.5,但消耗更少的 token 且运行成本更低。模型基于 1.5T 参数,定价更低、速度更快。xAI 因此进入前沿模型竞争行列。
Grok 4.5 lands fourth on the AA Intelligence ranki…
Grok 4.5 lands fourth on the AA Intelligence ranking with a score of 54 and is a clear third in the top AI labs, surpassing Google and Gemini.
It improves sixteen points from the last version and shows real strength in agentic work, which measures how well models handle practical tasks like knowledge jobs, tool use and decision making in realistic scenarios. On coding agent tests, which check performance as an actual coding assistant running through codebases and fixes, it matches top models like GPT-5.5 while using far fewer tokens and costing much less to run.
The efficiency stands out most for everyday use. Long agent runs or coding sessions stay affordable without sacrificing too much on results. Early testers report clearer improvements in design choices and self checking during real coding work, plus better handling of complex multi step tasks. It uses a much larger 1.5T param base and sits at a lower price point with solid speed.
This puts xAI firmly in the frontier conversation for practical developer and knowledge work, even if it does not top every single chart. The cost and token savings could make advanced agent tools more accessible for regular projects.
It's taken them a while to finally turn frontier. Good work, SpaceXAI. Hope it comes to lower price plans soon. No model card yet, though?
Remember, they have much bigger models (up to 10T) coming and a huge amount of GPUs to train them! https://t.co/Y1gyZxHttq