中国开源模型Kimi K3编码能力比Claude和GPT还强,价格却便宜好几倍,做长上下文编程选它准没错。
Kimi K3在Arena's Frontend Code盲测中击败Claude Fable 5和GPT-5.6 Sol,排名第一。在DeepSWE基准上,K3得69%且每任务成本$4.65,而Sol得73%成本$8.39,Fable 5得70%成本$21.63。在Artificial Analysis的Intelligence Index上,K3得57分,仅比Fable 5的60分低3分。该模型在长时编码和知识工作方面表现突出,在AA-Briefcase排名第二。但幻觉率为51%,接近Grok 4.5的54%。开放权重将于7月27日发布。
Kimi K3 took the top spot on Arena's Frontend Code…
Kimi K3 took the top spot on Arena's Frontend Code leaderboard days after launch, beating both Claude Fable 5 and GPT-5.6 Sol in blind human voting. On DeepSWE it scores 69% at $4.65 per task versus Sol's 73% at $8.39 and Fable 5's 70% at $21.63. That's a Chinese open-weight model competing on performance while undercutting on price.
Moonshot themselves are honest about where it sits. They say K3 trails Fable 5 and GPT-5.6 Sol overall but demonstrated frontier-level performance across their evaluation suite, consistently outperforming other tested models. On Artificial Analysis's Intelligence Index, it scores 57, only three points behind Fable 5 (60) and nearly tied with GPT-5.6 Sol high (56) and Opus 4.8 max (56). It's ahead of Grok 4.5 (54). That's extremely competitive with the frontier.
Where K3 stands out is long-horizon coding and knowledge work. It scored second on AA-Briefcase, Artificial Analysis's private benchmark of realistic knowledge work tasks, behind only Fable 5. In their kernel optimisation tests, it performed competitively with Fable 5 and substantially outperformed Opus 4.8, GPT-5.6 Sol and GPT-5.5. The model handles sustained multi-hour coding sessions, large repositories and vision-guided UI iteration unusually well.
On the Intelligence Index K3 costs $0.94 per task, roughly half of Opus 4.8 and slightly cheaper than Sol. The trade-off is token verbosity. K3 runs maximum reasoning effort by default and used 81k output tokens per DeepSWE task versus Sol's 60k, which means real task costs depend on how much the model decides to think out loud. Independent testing also shows a 51% hallucination rate, near Grok 4.5's 54%.
The open weights drop on 27 July. For now it works best as a specialist for long-context coding and research workflows where the price advantage matters and you can tolerate the verbosity. The fact that an open Chinese model can land #1 on a human-preference coding benchmark while scoring 57 on a broad independent index shows the gap is narrowing fast in specific areas.
https://t.co/uQexhZoQyx