行业71°

Arize AI 与 Fireworks AI 测试 10 模型,Kimi K3 接近 GPT-5.5

Many still debate open vs closed, and compare cost per token (accounting metric). Better to shift ...

精选理由

别只盯着 token 价格了,Arize 和 Fireworks 测了 10 个模型的 2400 次 agent 运行,发现按任务难度的路由能省钱又提高成功率。Kimi K3 还追平了 GPT-5.5。

AI 摘要

Arize AI 与 Fireworks AI 联合对 10 个模型进行了 2400 次 agent 运行基准测试,包括 Kimi K3。结果显示,Kimi K3 在整体表现上接近 GPT-5.5。不同模型在不同任务类型上各有优势,按任务难度路由可同时优化成本和覆盖率。重试和静默故障显著改变了经济性。开发者应关注每次成功任务的成本而非 token 价格。

原文 · Fireworks AI

Many still debate open vs closed, and compare cost per token (accounting metric). Better to shift ...

Many still debate open vs closed, and compare cost per token (accounting metric). Better to shift attention to cost per successful task. @seldo and the team at @arizeai did so across 2,400 runs. Conclusion: route by task difficulty, and you win on both cost and coverage. Arize AI @arizeai Token price tells you what a model costs to call. It does not tell you what it costs to finish the job. Arize and @FireworksAI_HQ benchmarked 10 models across 2,400 agent runs, including Kimi K3. In his tests, Arize's Head of DevRel @seldo found: • Kimi K3 nearly matched GPT-5.5 overall • Different models won on different task types • Routing improved cost and coverage • Retries and silent failures changed the economics completely Developers and product managers should optimize for cost per successful task, then use evals and traces to understand where each model belongs. Get the benchmark data: arize.com/blog/cost-per-… 🔗 View Quoted Tweet 💬 1 🔄 0 ❤️ 4 👀 757 📊 1 ⚡