阿里Qwen3.8-Flash-Next在Agent Arena排名第7,比前代模型提升2.4%,任务完成率高达12.3%。
Qwen3.8-Flash-Next由阿里巴巴发布,在Agent Arena中排名第7,在8.7K+真实智能体会话中实现+2.4%的净提升。该模型在任务完成确认成功指标上表现突出,达到+12.3%,在开放模型中排名第5。相比Qwen3.8-27B,Qwen3.8-Flash-Next在净提升和确认成功指标上均有显著优势。
Qwen3.8-Flash-Next by @Alibaba_Qwen has landed in Agent Arena, ranking #7 among open models (#24 ove...
Qwen3.8-Flash-Next by @Alibaba_Qwen has landed in Agent Arena, ranking #7 among open models ( #24 overall) with +2.4% net improvement across 8.7K+ real-world agentic sessions! Among open models, it sits just behind DeepSeek V4 Flash (High) at #6 (+3% net improvement) and two spots behind GLM-5.3 (Max) at #5 (+3.8%). By signal, Qwen3.8-Flash-Next stands out in delivering an explicit response from the community on task completion (Confirmed Success at +12.3%), ranking #5 among open models ( #7 overall). More detail on its performance by signal below. Qwen3.8-Flash-Next outperforms Qwen3.8-27B in overall net improvement (+2.4% vs. +1.5%) and Confirmed Success (+12.3% vs. +7.2%). Qwen3.8 Max remains ahead overall at +6%, with +10.8% Confirmed Success. Congrats to the @Alibaba_Qwen team on this contribution to the open ecosystem! Qwen @Alibaba_Qwen ⚡Meet Qwen3.8-Flash, a multimodal MoE and an early preview of the Qwen4 architecture, now open-weight! The production version Qwen3.8-Flash will be available soon via QwenCloud API at just $ 0.16/1M input tokens and $ 0.47/1M output tokens. 125B parameters + 51B N-gram embeddings, with just 6B activated per token. Unmatched cost-efficiency. What's new: 🥳 - Next architecture: GDN + QSA hybrid attention, Gated Residual, N-gram Embedding & Muon optimizer, serving as a precursor to the architecture used in Qwen4. - Dramatically lower training and inference costs: trained at just 1/9 the cost of Qwen3.7-Plus, while outperforming it across the board with especially strong gains in coding and office tasks. - Strong performance: scoring 58.7 on DeepSWE 1.1, 62.5 on SWE-bench Pro, 73.9 on CoWorkBench, 84.5 on AndroidWorld, and 95.7 on MathVision (with CI). - 262K native context, extensible to 1M with YaRN. We’re also releasing the weights for Qwen3.8-Flash-Next, giving the community an early look at the new architecture we’re exploring for Qwen4.🚀 We can't wait to see what you build with Qwen3.8-Flash!👀👇 - Bl qwen.ai/blog?id=qwen3.… FLgJ - Technical Repo github.com/QwenLM/Qwen3.8… IkQO - Hugging Fa huggingface.co/Qwen/Qwen3.8-F… AABt - ModelSco modelscope.cn/models/Qwen/Qw… NuFG 🔗 View Quoted Tweet 💬 2 🔄 6 ❤️ 34 👀 6012 📊 6 ⚡