Kimi K3:2.8T参数开源MoE模型,性能接近Claude Fable 5和GPT-5.6 Sol

Kimi K3: Open Frontier Intelligence

精选理由

Kimi 开源了2.8T参数的K3模型,激活104B参数,能处理百万token上下文,视觉能力也不弱,性能直逼Claude和GPT-5.6,适合用来部署或研究。

AI 摘要

Kimi K3是一个2.8万亿参数的Mixture-of-Experts模型,激活参数104B,支持100万token上下文窗口和原生视觉。它采用Kimi Delta Attention和Attention Residuals改进信息流动,训练效率相比Kimi K2提升约2.5倍。后训练中应用了强化学习,覆盖通用、智能体和编程领域,实现组合泛化。在长程编码、智能体、知识、推理和视觉任务上达到前沿水平,但整体性能仍落后于Claude Fable 5和GPT-5.6 Sol。官方已开源完整模型权重。

原文 · arXiv cs.LG

Kimi K3: Open Frontier Intelligence

We introduce Kimi K3, a 2.8T parameter Mixture-of-Experts model with 104 billion activated parameters, native vision capabilities, and a 1-million-token context window. Kimi K3 is built on Kimi Delta Attention and Attention Residuals, which improve information flow across sequence length and model depth. Together with Stable LatentMoE, which effectively activates 16 of 896 routed experts per token, and refined training and data recipes, these advances yield an approximately 2.5x improvement in overall scaling efficiency over Kimi K2. Post-training highlights reinforcement learning across general, agentic, and coding domains and multiple reasoning-effort levels, enabling compositional generalization and robust long-horizon execution. At 2.8T scale, Kimi K3 is supported by infrastructure advances in multiple areas: algorithm-system co-design for KDA, perfectly balanced expert-parallel training with efficient memory management, million-token agentic RL with persistent rollout and sandbox states, and deployment innovations. Extensive evaluations show that Kimi K3 achieves frontier-level performance across long-horizon coding, agentic, knowledge, reasoning, and vision tasks. While its overall performance still trails the most powerful proprietary models, namely Claude Fable 5 and GPT-5.6 Sol, Kimi K3 consistently outperforms other open and proprietary models evaluated in our suite. We release the full Kimi K3 model weights to facilitate future research and accelerate the broader deployment and adoption of frontier intelligence.

Kimi K3:2.8T参数开源MoE模型,性能接近Claude Fable 5和GPT-5.6 Sol · AI 热点