模型精选

DeepSeek V4-Flash 单 GPU 吞吐量对比引争议:72 TPS 仅为其三分之一

精选理由

有人实测推理只有 72 TPS,对比 DeepSeek V4-Flash 在 DSpark 报告里的单 GPU 成绩直接少了三分之二,评论里讨论了 TileLang 内核和 npugraph_ex 还能优化多少。

一条推文对比了 TileLang 内核下的推理吞吐量,实测约 72 TPS。作者指出这只有 DeepSeek V4-Flash 在 DSpark 报告中给出的单 GPU 吞吐量的 1/3。推文猜测可能是部署单元不够大,或者后续优化尚未跟上,并提到 npugraph_ex 的作用尚不确定。

原文 · Teortaxes

That's pretty bad? This is 72 TPS, and still the throughput is 1/3rd of what DeepSeek got "per GPU" with V4-Flash in DSpark report. this is with TileLang kernels. Maybe it'll be improved, or you need a bigger deployment unit. not sure what npugraph_ex will change https://t.co/D0LLYgFReA