模型多源确认91°

Grok 4.7 登陆 Agent Arena,同价较 4.6 有显著提升

Grok 4.7 by @SpaceXAI and @elonmusk is now in the Agent Arena! Your votes drive the @arena leaderbo...

精选理由

Grok 4.7 上线 Agent Arena 了,拿你最难的提示词去投票实测,看比 4.6 强多少。

Grok 4.7 现已加入 Agent Arena 排行榜,用户可提交自己的提示词参与投票。Agent Arena 基于全球用户贡献的数百万条真实长程智能体任务评估模型,参评模型可调用网络搜索、文件系统和终端工具完成复杂工作流。排行榜采用因果追踪(causal tracing)方法,衡量模型结果相对平均模型的表现。除 Agent Arena 外,Grok 4.7 还上线了文本、视觉、代码和文档四种 Battle Mode。官方称 Grok 4.7 在与 Grok 4.6 相同的价格和速度下有明显改进。

原文 · lmarena.ai

Grok 4.7 by @SpaceXAI and @elonmusk is now in the Agent Arena! Your votes drive the @arena leaderbo...

Grok 4.7 by @SpaceXAI and @elonmusk is now in the Agent Arena! Your votes drive the @arena leaderboards, head over and bring your toughest prompts. In Agent Arena, we measure models on millions of real-world, long-horizon agentic tasks from a global community of users. Models can access web search, filesystem, and terminal tools to complete complex workflows. The leaderboard measures model performance on outcomes relative to the average model using a causal tracing methodology. In addition to Agent Arena, Grok 4.7 is in Battle Mode for: Text, Vision, Code, and Document. SpaceXAI @SpaceXAI Grok 4.7 is here. It's a notable improvement over Grok 4.6 at the same price and speed. 🔗 View Quoted Tweet 💬 5 🔄 3 ❤️ 107 👀 9281 📊 12 ⚡