Alibaba Qwen 发布的 Qwen3.8-Flash-Next 在 Code Arena WebDev 获得优异成绩,性能提升显著,成本效率高,是 Qwen4 架构的早期预览,值得一看。
Alibaba Qwen 的 Qwen3.8-Flash-Next 在 Code Arena WebDev 获第 8 名,性能比 Qwen3.8-27B 提升 22 点,参数量 125B,成本效率高。新架构 GDN + QSA 混合注意力,训练和推理成本降低 1/9。即将通过 QwenCloud API 提供,价格 0.16/0.47 美元/1M tokens。
Exciting news: Qwen3.8-Flash-Next by @Alibaba_Qwen ranks ~#8 in Code Arena: WebDev (#3 among open mo...
Exciting news: Qwen3.8-Flash-Next by @Alibaba_Qwen ranks ~ #8 in Code Arena: WebDev ( #3 among open models) scoring 1617 (AutoEval). Priced at $0.16/$0.47 Mtokens, it reshapes the Code Arena: WebDev Pareto Frontier! This 125B parameter model (6B active) has been touted as a preview of the Qwen4 architecture, showing a +22pt stronger performance than Qwen3.8-27B. Note: this is an early AutoEval score, in which a Reward Model trained on Arena's human preference data casts automatic votes in place of live votes. We’ll continue to see how scores converge as more live human votes come in. See thread for more info on the methodology behind AutoEval. Congrats to the @Alibaba_Qwen team on the strong release! Qwen @Alibaba_Qwen ⚡Meet Qwen3.8-Flash, a multimodal MoE and an early preview of the Qwen4 architecture, now open-weight! The production version Qwen3.8-Flash will be available soon via QwenCloud API at just $ 0.16/1M input tokens and $ 0.47/1M output tokens. 125B parameters + 51B N-gram embeddings, with just 6B activated per token. Unmatched cost-efficiency. What's new: 🥳 - Next architecture: GDN + QSA hybrid attention, Gated Residual, N-gram Embedding & Muon optimizer, serving as a precursor to the architecture used in Qwen4. - Dramatically lower training and inference costs: trained at just 1/9 the cost of Qwen3.7-Plus, while outperforming it across the board with especially strong gains in coding and office tasks. - Strong performance: scoring 58.7 on DeepSWE 1.1, 62.5 on SWE-bench Pro, 73.9 on CoWorkBench, 84.5 on AndroidWorld, and 95.7 on MathVision (with CI). - 262K native context, extensible to 1M with YaRN. We’re also releasing the weights for Qwen3.8-Flash-Next, giving the community an early look at the new architecture we’re exploring for Qwen4.🚀 We can't wait to see what you build with Qwen3.8-Flash!👀👇 - Bl qwen.ai/blog?id=qwen3.… FLgJ - Technical Repo github.com/QwenLM/Qwen3.8… IkQO - Hugging Fa huggingface.co/Qwen/Qwen3.8-F… AABt - ModelSco modelscope.cn/models/Qwen/Qw… NuFG 🔗 View Quoted Tweet 💬 1 🔄 7 ❤️ 52 👀 5935 📊 7 ⚡