Qwen3.8-Flash-Next 开源,效率提升显著,成本降低,性能优异,是 Qwen4 架构的早期预览,值得尝试!
Qwen3.8-Flash-Next 作为 Qwen4 架构早期预览开源,效率提升至 Qwen3.7-Plus 的 1/3,训练成本降低至 1/9,在多个基准测试中表现优异。DeepSWE 1.1 得分 58.7,CoWorkBench 73.9。QwenCloud API 上提供,价格低至 0.16/1M 输入令牌和 0.47/1M 输出令牌。
Qwen3.8-Flash 正式发布, Qwen3.8-Flash-Next 正式开源! Qwen 官方把 Qwen3.8-Flash-Next 定义为 Qwen4 架构的早期预览,提前开源的目的,...
Qwen3.8-Flash 正式发布, Qwen3.8-Flash-Next 正式开源! Qwen 官方把 Qwen3.8-Flash-Next 定义为 Qwen4 架构的早期预览,提前开源的目的,是让 vLLM、SGLang 等推理框架和基础设施开发者在 Qwen4 正式家族发布前完成适配。 qwen.ai/blog?id=qwen3.… 效率方面:相对上一代旗舰 Qwen3.7-Plus(397B/A17B),Flash-Next 以 1/3 激活参数、1/3 训练 token、约 1/9 训练 FLOPs,在 14 个预训练基准上 8 个领先、其余最多落后 2.6 分。 稳定性方面:报告给出了可复现的压力测试方法(恒定学习率加倍模拟大规模失稳):在 4 倍最优学习率下,旧架构(AdamW + Qwen3.5 结构)频繁 loss 尖峰,而新配方(Muon + 门控残差)零尖峰;生产级训练全程未使用 qk-clip 之类的显式裁剪手段。 后训练模型的官方基准 对比对象为 Qwen3.7-Plus、DeepSeek-V4-Flash、Claude Opus 4.6 Max 等。 · 智能体编程:DeepSWE 1.1 得 58.7(Qwen3.7-Plus 仅 16.5,DeepSeek-V4-Flash 54.4);SWE-bench Pro 62.5;SWE-bench 多语言版 81.0 · 长程办公/Agent:CoWorkBench 73.9、JobBench 55.7、Toolathlon 73.5,均明显领先对比模型 · 通用:GPQA Diamond 91.7、LiveCodeBench v6 91.9、IFBench 81.3 · 多模态:AndroidWorld 84.5、MathVision 带代码解释器 95.7、OSWorld 2.0 部分 52.3、长视频 LVBench 76.6 Qwen @Alibaba_Qwen ⚡Meet Qwen3.8-Flash, a multimodal MoE and an early preview of the Qwen4 architecture, now open-weight! The production version Qwen3.8-Flash will be available soon via QwenCloud API at just $ 0.16/1M input tokens and $ 0.47/1M output tokens. 125B parameters + 51B N-gram embeddings, with just 6B activated per token. Unmatched cost-efficiency. What's new: 🥳 - Next architecture: GDN + QSA hybrid attention, Gated Residual, N-gram Embedding & Muon optimizer, serving as a precursor to the architecture used in Qwen4. - Dramatically lower training and inference costs: trained at just 1/9 the cost of Qwen3.7-Plus, while outperforming it across the board with especially strong gains in coding and office tasks. - Strong performance: scoring 58.7 on DeepSWE 1.1, 62.5 on SWE-bench Pro, 73.9 on CoWorkBench, 84.5 on AndroidWorld, and 95.7 on MathVision (with CI). - 262K native context, extensible to 1M with YaRN. We’re also releasing the weights for Qwen3.8-Flash-Next, giving the community an early look at the new architecture we’re exploring for Qwen4.🚀 We can't wait to see what you build with Qwen3.8-Flash!👀👇 - Bl qwen.ai/blog?id=qwen3.… FLgJ - Technical Repo github.com/QwenLM/Qwen3.8… IkQO - Hugging Fa huggingface.co/Qwen/Qwen3.8-F… AABt - ModelSco modelscope.cn/models/Qwen/Qw… NuFG 🔗 View Quoted Tweet 💬 1 🔄 0 ❤️ 2 👀 1113 📊 1 ⚡