SGLang在GB300部署DeepSeek-V4:5倍吞吐量提升

🚀 New blog: Serving DeepSeek-V4 on GB300 with SGLa…

精选理由

想用SGLang在GB300上榨干DeepSeek-V4?NVIDIA合作实测,吞吐翻5倍,交互延迟不变,MTP和量化细节全公开。

AI 摘要

与NVIDIA合作,在GB300上使用SGLang服务DeepSeek-V4,实现5倍吞吐量提升(~2,200→~11,200 tok/s/GPU,交互性~50 tok/s/user)。借助MTP,在80 tok/s/user交互性下吞吐再提升2.6倍。Blackwell Ultra聚合模式下30 tok/s/user时吞吐提升2.91倍,峰值无MTP吞吐提升超6倍。采用W4A4 MegaMoE量化(MXFP4)且精度损失可忽略。单个FP8-einsum修复将MTP接受率从0.57提至0.70。

原文 · LMSYS Org (SGLang)

🚀 New blog: Serving DeepSeek-V4 on GB300 with SGLa…

🚀 New blog: Serving DeepSeek-V4 on GB300 with SGLang: 5x Higher Throughput at the Same Interactivity Since Day-0

Together with @nvidia, we achieved 5X higher throughput at the same interactivity, serving DeepSeek-V4 on GB300 with SGLang.

Here's how the DeepSeek-V4 serving frontier moved on the public @SemiAnalysis_ InferenceX dashboard: 1️⃣ 5X throughput on GB300 disaggregated: ~2,200 → ~11,200 tok/s/GPU at ~50 tok/s/user 2️⃣ 2.6X more throughput at 80 tok/s/user with MTP. Curves now hold deep into the high-interactivity range deployments actually target 3️⃣ 2.91X on Blackwell Ultra aggregated at 30 tok/s/user, with 6X+ peak no-MTP throughput 4️⃣ W4A4 MegaMoE: activations now quantized to MXFP4 with negligible accuracy loss 5️⃣ A single FP8-einsum fix lifted MTP acceptance 0.57 → 0.70

Huge thanks to @NVIDIAAI @radixark for the deep collaboration on this! SGLang is PyTorch-native, and we're excited to share the full write-up on the @PyTorch blog!