蚂蚁的Ling-3.0-flash参数大但激活少,256K上下文,vLLM确认后续支持,开源在即。
蚂蚁集团Ling团队发布Ling-3.0-flash,一款124B参数的MoE模型,每个token仅激活5.1B参数。该模型采用混合线性注意力机制,支持原生256K上下文窗口,面向生产级智能体系统。vLLM项目祝贺其发布,并称赞其“先公告后开源”的发布模式。vLLM将在模型权重开源时提供支持,预计vLLM开源支持即将到来。
Congratulations to the Ling team @AntLingAGI on th…
Congratulations to the Ling team @AntLingAGI on the release of Ling-3.0-flash! 🎉 A 124B MoE model activating just 5.1B parameters per token—with hybrid-linear attention and native 256K context—is an exciting contribution to production-scale agent systems.
We also applaud the team’s announce-first, open-source-next release approach. Separating the announcement from the weight release lets the model team freeze the final checkpoint, configuration, tokenizer, and serving semantics, while giving open-source inference projects a stable window for correctness testing, performance tuning, Docker builds, and recipe validation. The community also gets a clear timeline instead of an ambiguous “coming soon.”
This is not a step back from day-0 support—it is a more sustainable way to deliver day-0 support for the artifacts users will actually run.
vLLM’s open-source support for Ling-3.0-flash is coming soon and will be available when the model weights are open-sourced. We hope more model vendors adopt this release pattern in the future!
Congrats again—we can’t wait to see what the community builds! 🚀