AI模型精选

Nvidia Groq 3 LPX 四倍于 Cerebras,但计算更复杂

Nvidia says its Groq 3 LPX is four times faster than Cerebras, but the math is more complicated

精选理由

Nvidia 的 Groq 3 LPX 芯片在性能上超越了 Cerebras,但需要更多加速器,了解其扩展性如何。

AI 摘要

Nvidia 的 Groq 3 LPX 推理芯片进入量产,Gemma 4 31B 上每秒处理 3,400 个 token,是 Cerebras 的四倍。但需要至少 64 个加速器,而 Cerebras 只需一个或两个,架构在大 MoE 模型上的扩展性仍待定。

原文 · Decoder

Nvidia says its Groq 3 LPX is four times faster than Cerebras, but the math is more complicated

Nvidia is moving its Groq 3 LPX inference chip into full production and reports 3,400 tokens per second on Gemma 4 31B, four times faster than Cerebras. But the numbers don't tell the whole story. Nvidia needs at least 64 accelerators to get there, while Cerebras needs only one or two, according to The Register. How well the architecture scales with large MoE models remains an open question. The article Nvidia says its Groq 3 LPX is four times faster than Cerebras, but the math is more complicated appeared first on The Decoder .