模型多源确认精选

SPLASH:大模型服务中切换注意力并行布局

SPLASH: Switching Parallel Layouts of Attention with Seamless Handoff for LLM Serving

精选理由

SPLASH让大模型服务能动态切换并行策略,在GLM和DeepSeek模型上显著提升吞吐,解决固定布局的局限性。

SPLASH系统可在请求运行时切换注意力并行布局,支持张量并行、数据并行、上下文并行和新型解耦所有权并行(DOP)。在B200 GPU上服务GLM-5.3时,SPLASH比固定布局部署提升1.3-1.73倍端到端吞吐量。DOP比数据并行注意力提供27-60%更多KV容量,且不复制数据。该系统在DeepSeek-V3.2(H200)和GLM-5.3-Flash(DCU)上也展现出相同布局优势。

原文 · arXiv: DeepSeek

SPLASH: Switching Parallel Layouts of Attention with Seamless Handoff for LLM Serving

No single way of parallelizing attention serves large language models well under all loads. Low concurrency favors tensor parallelism, many independent requests favor data-parallel attention, and long prompts favor context parallelism. Reasoning, agentic, and RL-rollout workloads make a fixed choice untenable: a batch that begins as many short requests ends as a few very long ones, so the best layout changes while the same requests run. Serving engines nevertheless fix one layout at launch, because changing it has meant draining requests and restarting workers. We present SPLASH, a serving system that switches the parallel layout of attention while requests are running. It builds on one observation: modern attention, with few or no KV heads, decouples where a request's KV cache lives from how attention weights are sharded. This has two consequences. First, layouts differ only in who owns the weights and the cache, and most of that state already sits where the next layout needs it; SPLASH reuses it, moves the rest in the background of ongoing inference, and hands off at a batch boundary, making a switch nearly free: its median overhead is under 0.51% of the step it runs in. Second, the decoupling exposes a layout that existing engines lack: Decoupled Ownership Parallelism (DOP) shards attention weights as tensor parallelism does while keeping each request's cache on a single owner as data-parallel attention does. DOP replicates neither, offers 27-60% more KV capacity than data-parallel attention, and gives the scheduler a choice when KV memory limits admission. A transition-aware scheduler follows the best of the four layouts as load changes. On B200 GPUs serving GLM-5.3, SPLASH improves end-to-end serving throughput by 1.3-1.73x over fixed-layout deployments, and the same layout regimes appear with DeepSeek-V3.2 on H200 and GLM-5.3-Flash on DCU.