Qwen3.8-Flash发布,多模态MoE模型新进展

What's better than an open-weight multimodal model release? Well, the technical report. I just love...

精选理由

想了解Qwen3.8-Flash的最新进展和成本效率,这是必读的技术报告。

AI 摘要

Qwen3.8-Flash是阿里巴巴Qwen团队发布的最新高效多模态MoE模型,具有125B参数和51B N-gram嵌入,每token激活6B。模型在DeepSWE、SWE-bench Pro、CoWorkBench等基准测试中表现出色,成本效率高。生产版本即将通过QwenCloud API提供,价格为0.16/1M输入token和0.47/1M输出token。

原文 · elvis

What's better than an open-weight multimodal model release? Well, the technical report. I just love...

What's better than an open-weight multimodal model release? Well, the technical report. I just love how these labs like Qwen and DeepSeek continue to drop gem after gem. Qwen3.8-Flash is the latest in efficient multimodal MoE models. Worth reading the report. Qwen @Alibaba_Qwen ⚡Meet Qwen3.8-Flash, a multimodal MoE and an early preview of the Qwen4 architecture, now open-weight! The production version Qwen3.8-Flash will be available soon via QwenCloud API at just $ 0.16/1M input tokens and $ 0.47/1M output tokens. 125B parameters + 51B N-gram embeddings, with just 6B activated per token. Unmatched cost-efficiency. What's new: 🥳 - Next architecture: GDN + QSA hybrid attention, Gated Residual, N-gram Embedding & Muon optimizer, serving as a precursor to the architecture used in Qwen4. - Dramatically lower training and inference costs: trained at just 1/9 the cost of Qwen3.7-Plus, while outperforming it across the board with especially strong gains in coding and office tasks. - Strong performance: scoring 58.7 on DeepSWE 1.1, 62.5 on SWE-bench Pro, 73.9 on CoWorkBench, 84.5 on AndroidWorld, and 95.7 on MathVision (with CI). - 262K native context, extensible to 1M with YaRN. We’re also releasing the weights for Qwen3.8-Flash-Next, giving the community an early look at the new architecture we’re exploring for Qwen4.🚀 We can't wait to see what you build with Qwen3.8-Flash!👀👇 - Bl qwen.ai/blog?id=qwen3.… FLgJ - Technical Repo github.com/QwenLM/Qwen3.8… IkQO - Hugging Fa huggingface.co/Qwen/Qwen3.8-F… AABt - ModelSco modelscope.cn/models/Qwen/Qw… NuFG 🔗 View Quoted Tweet 💬 3 🔄 3 ❤️ 14 👀 2773 📊 5 ⚡

Qwen3.8-Flash发布,多模态MoE模型新进展 · AI 热点