AI模型精选

NVIDIA Nemotron 3 Ultra 登上 Agent Arena 排行榜第20名

The newest open model to join the Agent Arena leaderboard, Nemotron 3 Ultra by @NVIDIA lands at #20 ...

精选理由

NVIDIA 开源模型在智能体评测中排第5

AI 摘要

NVIDIA 的 Nemotron 3 Ultra 在 Agent Arena 排行榜上位列第20名,在开源模型中排第5。该模型在用户表扬与投诉的净差值和工具幻觉率方面表现突出,但在可操控性和 bash 恢复能力上存在短板。排行榜基于30万+任务、200万+工具调用和4000万行代码的评测数据。当前分数置信区间较宽,排名仍在稳定中。

原文 · lmarena.ai

The newest open model to join the Agent Arena leaderboard, Nemotron 3 Ultra by @NVIDIA lands at #20 ...

The newest open model to join the Agent Arena leaderboard, Nemotron 3 Ultra by @NVIDIA lands at #20 overall and #5 among open models. Its standout signals are a positive praise-vs-complaint margin and low tool hallucination, but it's held back by steerability and bash recovery. Note the wide confidence intervals as scores are still stabilizing. In Agent Arena, models get web search, filesystem, and terminal tools to complete complex workflows: writing code, creating slide deck, researching the web, building apps, and analyzing documents. We use the causal tracing methodology to measure a model's net improvement which indicates how much it improves outcomes relative to the average model. See in thread how Nemotron 3 Ultra scored across 5 signals, drawn from tasks submitted by a global community of users. Arena.ai @arena Introducing Agent Arena: real-world agentic evals at scale. How do you evaluate agents doing actual work? We measure millions of live sessions where real users accomplish real tasks. On Arena, models now get web search, filesystem, and terminal tools to complete complex workflows: writing code, creating slide deck, researching the web, building apps, and analyzing documents. Every session produces rich signals. Users iterate with the agent turn-by-turn: approving, editing, correcting, praise or expressing frustration. The environment gives feedback too: shell errors, tool failures, recovery attempts, and more. Our leaderboard measures each model's agentic performance using causal inference across five signals: task success, steerability, error recovery, user praise vs. complaint, and tool hallucination. This leaderboard snapshot is built from 300K+ tasks, 2M+ tool calls, and 40M lines of code by agents. Top labs in Agent Arena: - #1 @OpenAI : GPT-5.5 (High) - #2 @AnthropicAI : Claude-Opus-4.7 (Thinking) - #3 @Zai_org : GLM-5.1 - #4 @GoogleDeepMind : Gemini-3.1-Pro - #5 @Kimi_Moonshot : Kimi-K2.6 More analysis in the thread, with the full technical blog below. 🔗 View Quoted Tweet 💬 3 🔄 1 ❤️ 47 👀 4311 📊 6 ⚡

  • Sebastian Raschka06-12 04:42原文
  • vLLM06-12 14:47原文
  • ollama06-13 01:26原文
  • Richard Socher06-11 15:30原文
  • NVIDIA AI06-11 18:04原文
  • Together AI06-11 20:04原文
  • karminski-牙医 (AI工具)06-12 04:31原文
  • LMSYS Org (SGLang)06-12 14:18原文
  • rohanpaul_ai06-13 01:55原文
  • Tri Dao (FlashAttention)06-12 04:20原文