Kimi K3 (Max) 登顶 Frontend Code Arena,开放权重和整体均第一

In Frontend Code Arena, Kimi K3 (Max) by @Kimi_Moonshot is ranked #1 among open and #1 overall! It’...

精选理由

Kimi K3 (Max) 在代码和智能体测试中双料夺冠,连 GLM-5.2 都比它低一截,开发者和研究者可以关注这个新标杆。

AI 摘要

Arena.ai 发布最新榜单,Kimi K3 (Max) 在 Frontend Code Arena 中以 1682 分排名开放权重第一,同时超越所有闭源模型成为整体第一。它在 Agent Arena 中取得 +9.75% 的净提升,领先 GLM-5.2 (Max) 的 +7.12%,并夺得开放权重组冠军。该模型在 7 个领域中的 5 个获得整体第一,仅在游戏和内容创作工具领域排名第二。Agent Arena 通过数百万真实长程智能体任务评估模型,涉及网页搜索、文件系统和终端工具。

原文 · lmarena.ai

In Frontend Code Arena, Kimi K3 (Max) by @Kimi_Moonshot is ranked #1 among open and #1 overall! It’...

In Frontend Code Arena, Kimi K3 (Max) by @Kimi_Moonshot is ranked #1 among open and #1 overall! It’s the #1 open-weight model in all Domains: Brand & Marketing, Reference-Based Design, Data & Analytics, Consumer Product, Gaming, Simulations, and Content Creation Tools. And for overall, #1 in 5 of 7 Domains, landing #2 only in Gaming and Content Creation Tools. Arena.ai @arena Big update: Among open-weight models, Kimi K3 (Max) is #1 in the Agent Arena with +9.75% net-improvement, surpassing GLM-5.2 (Max) at +7.12%, and landed the #1 spot across 5 signals (see below). Kimi K3 (Max) is also now #1 in open-weight in the Frontend Code (1682 pts) and Text (1485 pts) Arenas. Agent Arena measures models on millions of real-world, long-horizon agentic tasks. Models get web search, filesystem, and terminal tools to complete complex workflows: writing code, creating slide decks, researching the web, building apps, and analyzing documents. We use causal tracing methodology to measure a model's net improvement, which indicates how much it improves outcomes relative to the average model. Congrats to the @Kimi_Moonshot team for their contribution to the open ecosystem. 🔗 View Quoted Tweet 💬 7 🔄 5 ❤️ 32 👀 2256 📊 8 ⚡

Kimi K3 (Max) 登顶 Frontend Code Arena,开放权重和整体均第一 · AI 热点