模型多源确认精选73°

Fireworks 发布 Ember-1:基于 Kimi K3,token 用量省近四成

精选理由

Fireworks 在 Kimi K3 上做了后训练,产出 Ember-1,跑编程任务时推理 token 少了 71%,性能不变,省钱明显。

Fireworks Research 团队推出 Ember-1,基于 Kimi K3 做后训练,让模型减少重复推理。推理模型通常把 90% 以上的输出 token 花在思考上,在智能体循环里会反复思考同样的问题。Fireworks RL 用真实智能体编程任务循环做训练,教会模型区分哪些推理能改变答案、哪些只是原地打转。在线 A/B 测试中,Ember-1 在相同成功率下推理 token 减少 71%,总 token 减少 39%。

原文 · Cline

Ember-1 by the Fireworks Research team is built on Kimi K3, and uses ~40% fewer tokens while achieving the same performance on benchmarks.

This was accomplished by post-training K3 to think less repetitively.

Reasoning models spend most of their output tokens (sometimes 90%+) on thinking before they answer, which gets expensive in agentic loops where the model tends to re-think the same thoughts on every step.

Fireworks RL trained on real agentic coding task loops to teach the model which reasoning actually changes the answer vs which is just looping.

In a live A/B test on coding traffic, Ember-1 used 71% fewer reasoning tokens and 39% fewer total tokens than K3 at the same success rate.