@vllm_project让Qwen3.8-Flash-Next在NVIDIA和AMD上无缝运行,新模型带来更多可能性,值得一看!
@vllm_project使Qwen3.8-Flash-Next从第一天起就在NVIDIA和AMD上运行,超稀疏多模态MoE模型,125B参数,6B活跃,262K本地,1M通过YaRN,额外51B N-gram表可卸载,Qwen Sparse Attention是新的引擎工作所在。
Amazing! 🥳 Thanks @vllm_project for getting Qwen3.8-Flash-Next running on NVIDIA and AMD from day 0...
Amazing! 🥳 Thanks @vllm_project for getting Qwen3.8-Flash-Next running on NVIDIA and AMD from day 0. vLLM @vllm_project Qwen3.8-Flash-Next from @Alibaba_Qwen has day-0 support in vLLM, verified on NVIDIA and AMD GPUs. 🎉 Ultra-sparse multimodal MoE: 125B params, 6B active, 262K native, 1M via YaRN. On top of those sits a separate 51B N-gram table you can offload. Most of it will look familiar. The Gated DeltaNet layers reuse the KV path vLLM has had since Qwen3-Next: only a quarter of the layers hold a growing KV cache. Keep the 51B table in host RAM instead of HBM with VLLM_PLE_CPU_OFFLOAD=1. Qwen Sparse Attention is where the new engine work went. For now the model runs from vllm/vllm-openai:qwen38-flash-next. Thanks to @Alibaba_Qwen for the weights, and for opening them this early! 🙌 recipes.vllm.ai/Qwen/Qwen3.8-F… rGA 🔗 View Quoted Tweet 💬 14 🔄 10 ❤️ 301 👀 21501 📊 33 ⚡