AMD 将在 PyTorch 大会展示 MoRI + vLLM 跨 pod 的 MoE 推理方案
AMD 把 MoRI 通信栈接进 vLLM,让 MoE 推理能跨 pod 用 RDMA 分摊专家和搬 KV cache,搞大模型部署的可以看看基准数据。
AMD 团队 Ravi Gupta、Rishi Madduri 和 Shiksha Patel 将在 PyTorch Conference North America 2026 上介绍 MoRI + vLLM 项目。该方案把 MoRI 通信栈集成进 vLLM,通过 RDMA 实现 Wide Expert Parallelism,让 MoE 模型的专家可以分布在多个 pod 上,并在 AMD Instinct GPU 间传输 KV cache,实现 prefill 与 decode 分离部署。演讲还会覆盖跨 pod 时的正确性与稳定性问题,以及长上下文、高并发负载下的 vLLM 基准结果。
Running large Mixture-of-Experts models across pods creates two major challenges: expert dispatch and combine traffic, and moving the KV cache between prefill and decode.
At #PyTorchCon North America 2026, Ravi Gupta, Rishi Madduri, and Shiksha Patel (@AMD) will present “MoRI + vLLM: Wide Expert Parallelism and RDMA KV-Cache Transfer for Disaggregated MoE Serving on AMD.”
The poster will show how bringing the MoRI communication stack into @vllm_project enables Wide Expert Parallelism across pods using RDMA and moves the KV cache for disaggregated prefill and decode on AMD Instinct GPUs.
They will also cover correctness and stability issues encountered when spanning pods and share vLLM benchmark results for long, high-concurrency workloads.
Register for PyTorch Conference North America 2026: https://t.co/jBApW8nESi