腾讯混元Hy3在Agent Arena和代码榜单上表现不错,工具使用能力尤其突出,适合关注开源智能体模型的朋友看。
腾讯混元发布Hy3模型,在Arena.ai的Agent Arena中位列总榜第25名,开源权重模型中排第5。在Frontend Code Arena中排名开源模型第2、总榜第16。Agent Arena基于百万级真实长周期智能体任务评测,Hy3在工具使用(CLI/bash错误恢复)上净提升+2.6%,排名第25;可操纵性是其最大弱点,净提升-7.1%,排名第30。
Agent & Coding 🔥🔥🔥
Agent & Coding 🔥🔥🔥 Arena.ai @arena Hy3 by Tencent is #5 in Agent Arena for open-weight models ( #25 overall)! It also ranks as the #2 open model in the Frontend Code Arena ( #16 overall)! In Agent Arena: Hy3 lands at #25 overall (net -2.2%). Hy3 has strengths in tool-use (recovering well from CLI/bash errors, +2.6% and #25 ), its biggest weakness is steerability as it struggles to course-correct when users push back, coming in at -7.1% ( #30 ). Agent Arena measures models on millions of real-world, long-horizon agentic tasks. Models get web search, filesystem, and terminal tools to complete complex workflows: writing code, creating slide decks, researching the web, building apps, and analyzing documents. We use causal tracing methodology to measure a model's net improvement, which indicates how much it improves outcomes relative to the average model. Below we break down how Hy3 scored across 5 key signals, drawn from tasks submitted by a global community of users. Here’s an overview on the signals: User-satisfaction proxies - Confirmed Success: an explicit "yes that worked" feedback from the user - Praise vs. Complaint: implicit sentiment in users reactions - Steerability: can the model course-correct when you push back? Tool-use proxies - Bash Recovery: how it recovers from CLI errors (primary signal for tool use) - Tool Hallucination: does it call tools that don't exist Congrats to the @TencentHunyuan team on this release! 🔗 View Quoted Tweet 💬 3 🔄 0 ❤️ 26 👀 1788 📊 5 ⚡