技巧精选

Kimi K3实测:2.8T参数、1M上下文,claude-context省40% token

When 𝗞𝗶𝗺𝗶 𝗞𝟯 came out, we tested it right away. This is our experimental sharing: Kimi K3 has...

精选理由

实测Kimi K3确实强,但烧token也快。作者分享了用claude-context省40% token的实操,适合搞智能体的团队参考。

AI 摘要

Kimi K3拥有2.8T参数和1M token上下文窗口,支持原生的视觉理解,在编码智能体基准上得分很高。实际测试中,营销团队用K3一个下午就完成了活动登陆页原型。但token开销巨大:cache-hit输入0.30美元/MTok,cache-miss输入3.00美元/MTok,输出15.00美元/MTok。通过添加claude-context MCP工具,实现语义代码搜索,在保持同等检索质量下减少约40%的token消耗。

原文 · Milvus

When 𝗞𝗶𝗺𝗶 𝗞𝟯 came out, we tested it right away. This is our experimental sharing: Kimi K3 has...

When 𝗞𝗶𝗺𝗶 𝗞𝟯 came out, we tested it right away. This is our experimental sharing: Kimi K3 has 2.8T parameters, a 1M-token context window, native visual understanding, and strong coding-agent scores. After a few internal runs, we decided to move part of our workflow from Fable to K3. The first surprise came from marketing. One teammate in marketing used K3 to prototype a campaign landing page in one afternoon. Not production-ready, but enough to align the idea before involving engineering. That was the good part. The less fun part was the token bill. Large-context coding agents can spend a lot of tokens just figuring out where things are in a repo. Kimi's pricing makes that visible: cache-hit input is $0.30/MTok, cache-miss input is $3.00/MTok, and output is $15.00/MTok. So we added 𝗰𝗹𝗮𝘂𝗱𝗲-𝗰𝗼𝗻𝘁𝗲𝘅𝘁. It gives coding agents semantic code search, so they retrieve relevant files and snippets instead of loading whole directories into the prompt. In our controlled evaluation, Claude Context MCP reduced tokens by around 40% at equivalent retrieval quality. K3 made the agent stronger. claude-context made it cheaper to aim that strength at the ri github.com/zilliztech/cla… s://t.co/yyWCfHbgZv 💬 0 🔄 0 ❤️ 0 👀 54 ⚡