做企业级 AI 推理或编程智能体的团队,如果被 GPU 集群的延迟和带宽瓶颈困扰,Cerebras 的晶圆级方案值得关注——它用硬件架构创新解决了模型权重和激活值传输的痛点,实测数据比 GPU 云快一个数量级。
Cerebras 宣布其晶圆级芯片在 1 万亿参数的 Kimi K2.6 模型上达到了 981 tokens/sec 的推理速度,经 Artificial Analysis 验证,比最快的 GPU 云快 6.7 倍。传统 GPU 集群因跨芯片拆分模型导致大量数据传递延迟,而 Cerebras 的晶圆级芯片将整个处理器构建在单个硅晶圆上,片上路由带宽更高、延迟更低。这一速度优势对于企业级编程智能体等需要快速迭代测试和调试的场景尤为关键。Cerebras 声称其真正的商业价值不在于单纯的速度,而在于能在足够大的模型上实现这种速度,从而支撑企业级应用。
Cerebras reported 981 tokens/sec on the 1T-paramet…
Cerebras reported 981 tokens/sec on the 1T-parameter Kimi K2.6 model. 6.7× faster than the next GPU cloud, validated by Artificial Analysis.
The hard part is moving model weights and activations fast enough, because normal GPU clusters split the model across many chips and spend a lot of time passing data between them.
Cerebras uses wafer-scale chips, meaning one processor is built across a full silicon wafer, so more of the routing happens on-chip with much higher bandwidth and lower delay.
The real business claim is not just speed, but speed on a model big enough for enterprise coding agents, where every extra second slows testing, debugging, and iteration.
---
cerebras. ai/blog/cerebras-kimi-k2-Enterprise