AI模型精选

MiniMax M3 在 Agent Arena 排名第18,成为第5开源模型

MiniMax M3 ranks #18 on the new Agent Arena leaderboard, and is now the #5 open model! The open-wei...

精选理由

MiniMax M3 在 Agent Arena 上排名上升了4位,是最强开源模型之一,能写代码、做PPT、查资料,幻觉控制也顶级。

AI 摘要

MiniMax M3 在全新 Agent Arena 排行榜上位列第18,是排名第5的开源模型。相比 M2.7,M3 从第22名升至第18名,主要改进是任务成功确认和 bash 错误恢复能力。工具幻觉保持低位,与最佳模型并列第一。排行榜基于30万+任务、200万+工具调用和4000万行代码的代理会话评估。

原文 · lmarena.ai

MiniMax M3 ranks #18 on the new Agent Arena leaderboard, and is now the #5 open model! The open-wei...

MiniMax M3 ranks #18 on the new Agent Arena leaderboard, and is now the #5 open model! The open-weight model improves meaningfully over MiniMax M2.7, climbing from #22 to #18 . Its biggest gain is confirmed task success, and recovering from bash errors. Tool hallucination stays low across both versions, tied for first. In Agent Arena, we measure models on millions of real-world, long-horizon agentic tasks and then use a causal tracing methodology to measure a model's net improvement which indicates how much it improves outcomes relative to the average model. Below we break down how MiniMax M3 from @MiniMax_AI scored across 5 signals, drawn from tasks submitted by a global community of users. Arena.ai @arena Introducing Agent Arena: real-world agentic evals at scale. How do you evaluate agents doing actual work? We measure millions of live sessions where real users accomplish real tasks. On Arena, models now get web search, filesystem, and terminal tools to complete complex workflows: writing code, creating slide deck, researching the web, building apps, and analyzing documents. Every session produces rich signals. Users iterate with the agent turn-by-turn: approving, editing, correcting, praise or expressing frustration. The environment gives feedback too: shell errors, tool failures, recovery attempts, and more. Our leaderboard measures each model's agentic performance using causal inference across five signals: task success, steerability, error recovery, user praise vs. complaint, and tool hallucination. This leaderboard snapshot is built from 300K+ tasks, 2M+ tool calls, and 40M lines of code by agents. Top labs in Agent Arena: - #1 @OpenAI : GPT-5.5 (High) - #2 @AnthropicAI : Claude-Opus-4.7 (Thinking) - #3 @Zai_org : GLM-5.1 - #4 @GoogleDeepMind : Gemini-3.1-Pro - #5 @Kimi_Moonshot : Kimi-K2.6 More analysis in the thread, with the full technical blog below. 🔗 View Quoted Tweet 💬 9 🔄 4 ❤️ 93 👀 7821 📊 14 ⚡