模型75°

Grok 4.7 发布,上线 Agent Arena 接受真实智能体任务评测

Grok 4.7 by @SpaceXAI just dropped. Looking at how past Grok versions have trended on Agent Arena's ...

精选理由

Grok 4.7 上 Agent Arena 了,拿真实智能体任务投票打分,想看排名可以去投一票。

Grok 4.7 已发布并进入 Agent Arena 排行榜,用户可以用自己最难的 prompt 投票,得分随投票陆续公布。Agent Arena 基于来自全球用户的数百万条真实长程智能体任务评测模型,允许模型调用网页搜索、文件系统和终端工具完成复杂工作流。排行榜用 causal tracing 方法衡量模型结果相对平均模型的表现,官方还根据历代 Grok 在 net improvement score 上的走势发起投票猜 4.7 的落点。除 Agent Arena 外,Grok 4.7 已进入 Text、Vision、Code、Document 四个 Battle Mode 类别。

原文 · lmarena.ai

Grok 4.7 by @SpaceXAI just dropped. Looking at how past Grok versions have trended on Agent Arena's ...

Grok 4.7 by @SpaceXAI just dropped. Looking at how past Grok versions have trended on Agent Arena's net improvement score, where do you think 4.7 lands? Poll below. Scores coming soon as you vote on @arena ! Arena.ai @arena Grok 4.7 by @SpaceXAI and @elonmusk is now in the Agent Arena! Your votes drive the @arena leaderboards, head over and bring your toughest prompts. In Agent Arena, we measure models on millions of real-world, long-horizon agentic tasks from a global community of users. Models can access web search, filesystem, and terminal tools to complete complex workflows. The leaderboard measures model performance on outcomes relative to the average model using a causal tracing methodology. In addition to Agent Arena, Grok 4.7 is in Battle Mode for: Text, Vision, Code, and Document. 🔗 View Quoted Tweet 💬 1 🔄 5 ❤️ 37 👀 3867 📊 5 ⚡