想了解阿里巴巴如何以更低成本打造强大模型?Qwen3.8-Flash-Next值得一试,它比DeepSeek-V4-Flash和Claude Opus 4.6更高效,成本更低。
阿里巴巴Qwen团队发布Qwen3.8-Flash-Next,混合专家模型,每token激活125亿参数中的6个。训练成本降低至九分之一,在编码和办公基准测试中击败DeepSeek-V4-Flash和Claude Opus 4.6,对OpenAI和Anthropic构成更多价格压力。
Alibaba releases Qwen3.8-Flash-Next, targeting "ultimate cost efficiency"
Alibaba's Qwen team is previewing the Qwen4 architecture with Qwen3.8-Flash-Next, a mixture-of-experts model that activates just 6 out of 125 billion parameters per token. At one-ninth the training cost, it beats much larger competitors like DeepSeek-V4-Flash and Claude Opus 4.6 on coding and office benchmarks, adding more pricing pressure on OpenAI and Anthropic. The article Alibaba releases Qwen3.8-Flash-Next, targeting "ultimate cost efficiency" appeared first on The Decoder .