MegaFlux:通过流水线专家复制实现抗偏斜MoE大内核
MegaFlux: Skew-Resilient MoE Megakernels via Pipelined Expert Replication
清华团队推出MegaFlux,解决MoE模型GPU负载不均衡问题,在DeepSeek-V4-Pro上实测显著加速。
MegaFlux解决了MoE大内核中的路由偏斜问题,在8块NVIDIA B200 GPU上实现前向加速1.45倍,后向加速1.28倍。该技术通过流水线专家复制隐藏了56-76%的副本权重传输成本,在vLLM中应用于DeepSeek-V4-Pro预填充时获得1.13-1.26倍端到端加速。峰值性能分别达到2.14倍和2.64倍提升。
MegaFlux: Skew-Resilient MoE Megakernels via Pipelined Expert Replication
Mixture-of-experts (MoE) megakernels fuse expert-parallel communication with expert computation. However, under fixed expert placement, routing skew creates GPU stragglers: overloaded GPUs determine layer latency while others sit idle. Replicating hot experts can shift work to underloaded GPUs, but dynamic replicas introduce additional work: replicas must receive expert weights to execute and, during training, their partial weight gradients must be reduced at the expert owners. We present MegaFlux, which makes expert replication a runtime decision and pipelines the communication induced by replication within persistent MoE execution. An on-device planner jointly selects replica locations and assigns tile-aligned token blocks under a per-GPU replica budget, leaving router outputs unchanged. The forward and backward megakernels realize pipelined expert replication: replicas begin computation as their required weights arrive, while backward overlaps replica-gradient reduction with ongoing expert computation. MegaFlux extends TensorRT-LLM's CuTeDSL MegaMoE forward kernel and introduces a new backward MoE megakernel. Across 147 configurations per direction on eight NVIDIA B200 GPUs, MegaFlux achieves geometric-mean speedups of $1.45\times$ for forward and $1.28\times$ for backward over the same megakernels with fixed placement, peaking at $2.14\times$ and $2.64\times$. In ablations, pipelining hides $56$--$76$% of replica-weight transfer cost in forward and $91$--$100$% of combined weight-transfer and replica-gradient-reduction cost in backward, yielding up to $13.2$% and $26.7$% additional layer-latency reductions over the same replication plans with these operations executed separately. Integrated into vLLM for DeepSeek-V4-Pro prefill, MegaFlux delivers $1.13$--$1.26\times$ median end-to-end speedups over fixed placement.