阿里Qwen团队发布了Qwen3.8-Flash-Next模型,参数量巨大,架构创新,训练成本低,值得关注。
阿里Qwen团队发布Qwen3.8-Flash-Next模型,包含125B参数,6B激活参数。模型采用Gated DeltaNet、Qwen Sparse Attention、Gated Residual、N-gram Embedding等架构,训练成本仅为Qwen3.7-Plus的1/9。预览了Qwen4架构。
Alibaba’s Qwen Team Releases Qwen3.8-Flash-Next: A 125B Multimodal MoE With 6B Active Parameters Previewing the Qwen4 Architecture
We look at Qwen3.8-Flash-Next, Alibaba's open-weight multimodal Mixture-of-Experts model and an early preview of the Qwen4 architecture. We break down where the 180B parameters actually sit: a 125B backbone, a 51B N-gram embedding table, and a 4B multi-token prediction module, with only 6B active per token. We walk through the four architectural changes — the Gated DeltaNet and Qwen Sparse Attention hybrid, Gated Residual, N-gram Embedding, and the Muon optimizer. We also cover the benchmark results, the reported 1/9 training cost against Qwen3.7-Plus, and what self-hosting a 172.78 GiB FP8 checkpoint really demands. The post Alibaba’s Qwen Team Releases Qwen3.8-Flash-Next: A 125B Multimodal MoE With 6B Active Parameters Previewing the Qwen4 Architecture appeared first on MarkTechPost .