SpaceXAI 新发的 Grok 4.5 专门为编程和智能体任务优化,已经在 Agent Arena 开测,你可以亲自去给它投票。
Grok 4.5 由 SpaceXAI 发布,是首款针对编码和智能体任务训练的模型,并已进入 Agent Arena 基准测试。Agent Arena 基于百万级真实世界长时任务,允许模型使用网页搜索、文件系统和终端工具。排行榜采用因果追踪方法衡量模型相对于平均模型的结果表现。此外,Grok 4.5 还参与了 Battle Mode 中的文本、视觉和前端代码竞技场。
Grok 4.5 by @SpaceXAI is now in the Agent Arena! In Agent Arena, we measure models on millions of r...
Grok 4.5 by @SpaceXAI is now in the Agent Arena! In Agent Arena, we measure models on millions of real-world, long-horizon agentic tasks from a global community of users. Models can access web search, filesystem, and terminal tools to complete complex workflows. The leaderboard measures model performance on outcomes relative to the average model using a causal tracing methodology. In addition to Agent Arena, Grok 4.5 is in Battle Mode for: Text, Vision and Code Arena: Frontend. Come bring your toughest prompts - your votes drive the leaderboards. SpaceXAI @SpaceXAI Announcing Grok 4.5, our first model trained specifically for coding and agents. It was trained with Cursor and offers frontier intelligence at leading speeds and cost efficiency. x.ai/news/grok-4-5 Your browser does not support the video tag. 🔗 View on Twitter 🔗 View Quoted Tweet 💬 4 🔄 4 ❤️ 40 👀 4010 📊 7 ⚡