论文

CascadeEP:异步专家执行解决 MoE 推理注意力不均衡问题

CascadeEP: Asynchronous Expert Execution for MoE Prefill under Attention Imbalance

精选理由

MoE 推理部署的人可以看看,DeepSeek-V4 和 GLM-5.3 上实测 TTFT 快了近一半,思路是让专家计算别干等慢副本。

论文针对 MoE 推理中数据与专家并行(DEP)架构的调度问题:同步 EP 会等所有注意力副本的 token 到齐才开始专家 FFN 计算,造成等待。CascadeEP 提出三项机制:异步 EP 允许专家计算提前启动、streamFFN 对就绪 token 批处理、OEWF 让快的副本抓取专家权重执行其他 GPU 上未开始的工作。在 DeepSeek-V4-Flash、DeepSeek-V4-Pro 和 GLM-5.3 上的评测显示,p95 TTFT 最高提速 1.48 倍,推理吞吐最高提升 1.17 倍。

原文 · arXiv: DeepSeek

CascadeEP: Asynchronous Expert Execution for MoE Prefill under Attention Imbalance

Mixture-of-experts (MoE) serving commonly deploys data and expert parallelism (DEP): attention replicas run distinct request batches while routed experts are sharded across an expert-parallel (EP) group. During prefill, attention replicas finish dispatch at different times, but synchronous EP delays expert feed-forward network (FFN) computation until routed inputs from all replicas are ready. Request schedulers seek to balance load while reusing the key-value (KV) cache of shared prompt prefixes to avoid redundant prefill computation. These goals can conflict when a replica holding a matching prefix is already overloaded, leaving residual attention imbalance. We present ASYNCEP, a distributed execution engine for MoE prefill. ASYNCEP proposes three mechanisms. Asynchronous EP allows expert computation to start before tokens from all attention replicas are ready. streamFFN batches ready tokens to balance early execution with FFN computation efficiency. Opportunistic expert weight fetching (OEWF) allows a faster replica to fetch expert weights and execute unstarted work from other GPUs. We evaluate ASYNCEP on DeepSeek-V4-Flash, DeepSeek-V4-Pro, and GLM-5.3, and our results show that ASYNCEP achieves up to 1.48x speedup in p95 time-to-first-token (TTFT) and improves the inference throughput by up to 1.17x.