vLLM v0.31.0 发布
vLLM 团队发布了新版本,优化了多种模型支持和服务性能,特别适合需要大规模部署 AI 应用的开发者。
vLLM v0.31.0 版本发布,包含 717 次提交和 307 位贡献者。新版本支持 DeepSeek-V4.1-Flash 模型的 FlashMLA mega 注意力机制,以及 NVFP4 KV 缓存技术。引入了快速重启功能,可在重启时保持权重在 GPU 内存中,并支持数据并行和 MTP 草稿功能。
717 commits. 307 contributors. 96 first-timers. vLLM v0.31.0 is live. 🎉
Highlights:
🤖 DeepSeek-V4.1-Flash: FlashMLA mega attention with NVFP4 KV caching, DeepGEMM sparse MQA, Mega-Gate, and shared Engram host tables
⚡ Fast restart: vllm preload keeps weights in GPU memory across restarts, with data parallelism and MTP draft support; experimental engine snapshots
🛠️ Model Runner V2: draft-model speculative decoding, LiLiCorr, async DFlash, and DSpark adaptive verification for Gemma4
🌐 Large-scale serving: MoonEP balanced all-to-all, prefill context parallelism with data parallelism, DeepEPv2, and RL weight transfer
🎛️ Smarter scheduling: independent active-sequence limits, adaptive long-prefill chunking, and priority for requests already holding KV blocks
🗄️ HiSparse hardening: fixes for MTP verification, chunked-prefill livelocks, and host/GPU prefix caching
Thanks to everyone who made this release possible! 💙
Full release notes 👇 https://t.co/DJxJlbz7WI
- DAIR.AI02:00原文