Qwen3.8-27B发布:单GPU跑1M上下文,vLLM Day-0支持

One GPU, 1M context, Day-0 ready. Big props to the vLLM team for the seamless integration!👍 Try Qwe...

精选理由

阿里发了能塞进单卡GPU的27B模型,长上下文到1M,vLLM当天就能用。适合想自己部署长上下文开源模型的开发者。

AI 摘要

阿里发布 Qwen3.8-27B,单张 Blackwell GPU 即可部署,原生上下文 262K,可扩展至 1M。模型采用与 2.4T 旗舰相同的混合骨干架构,但为 dense 而非 MoE。vLLM 提供 Day-0 支持,内置 MTP 草稿头,短提示词接受率在 BF16 下为 92.2%、FP8 下为 84.8%。在 NVIDIA GB300 上完成端到端验证,支持 BF16/FP8 TP4 与 NVFP4 TP1,工具调用和 1M 上下文生成均正常。

原文 · 阿里通义 Qwen

One GPU, 1M context, Day-0 ready. Big props to the vLLM team for the seamless integration!👍 Try Qwe...

One GPU, 1M context, Day-0 ready. Big props to the vLLM team for the seamless integration!👍 Try Qwen3.8-27B on vLLM: @vllm_project recipes.vllm.ai/Qwen/Qwen3.8-2… c vLLM @vllm_project 🎉 Qwen3.8-27B is here from @Alibaba_Qwen , and the whole thing fits on a single GPU. Same hybrid backbone as the 2.4T flagship, dense instead of MoE. Day-0 support in vLLM. 🚀 What is in it for serving ✨ - Fits one Blackwell GPU in every precision. Qwen ships BF16 and FP8, the NVFP4 build from @inferact - 262K native context, stretching to 1M. At that length one GB300 still has room for roughly 6.6M KV tokens. Six full-length sequences in flight, on one GPU - An MTP draft head rides inside the checkpoint, so speculative decoding needs no separate speculator repo. Acceptance on short prompts measured 92.2% in BF16 and 84.8% in FP8 Verified end-to-end on @NVIDIA GB300: BF16 and FP8 at TP4, NVFP4 at TP1, tool calls working, correct generations at 1M context. Two prerequisites, vLLM nightly and transformers 5.8.0+. recipes.vllm.ai/Qwen/Qwen3.8-2… 93x 🔗 View Quoted Tweet 💬 5 🔄 10 ❤️ 145 👀 6999 📊 17 ⚡

Qwen3.8-27B发布:单GPU跑1M上下文,vLLM Day-0支持 · AI 热点