AI模型精选

Qwen3.8-Flash-Next 与 SGLang 部署

Big thanks to @sgl_project for the day-0 support! 🙌 Qwen3.8-Flash-Next is ready to deploy with SGLa...

精选理由

Alibaba_Qwen 发布 Qwen3.8-Flash-Next,与 SGLang 合作,模型容量大,嵌入量大,支持长任务,值得期待。

AI 摘要

Qwen3.8-Flash-Next 与 SGLang 部署,包含 125B 主模型和 51B N-gram 嵌入,GDN + QSA 混合注意力机制,Gated Residual 结构,Muon 训练。SGLang 为 Qwen4 新架构提供支持。

原文 · 阿里通义 Qwen

Big thanks to @sgl_project for the day-0 support! 🙌 Qwen3.8-Flash-Next is ready to deploy with SGLa...

Big thanks to @sgl_project for the day-0 support! 🙌 Qwen3.8-Flash-Next is ready to deploy with SGLang today. SGLang @sgl_project Congrats to @Alibaba_Qwen on launching Qwen3.8-Flash! SGLang is proud to be a day-0 partner supporting the new architecture preview for Qwen4. It's a 125B main model with 51B of N-gram embeddings and 6B activated per token. The 51B N-gram embeddings scale model capacity with almost no extra compute per token, and can sit in host memory with async prefetch instead of occupying GPU memory. The GDN + QSA hybrid attention gives you efficient memory and precise retrieval at the same time on long-horizon tasks, while Gated Residual gives the model 4 lanes instead of 1 to pass information between layers. And it's trained with Muon! We're excited for what's next with Qwen4, and we already have plenty of ideas for how to use the N-gram embeddings in new deployment setups. Stay tuned! Blog and cookbook in the comments👇 🔗 View Quoted Tweet 💬 7 🔄 2 ❤️ 120 👀 11800 📊 13 ⚡

Qwen3.8-Flash-Next 与 SGLang 部署 · AI 热点