IQuest-Q1 发布:320B MoE 编程模型,vLLM 首日支持
IQuest 新出的 320B 编程模型 IQuest-Q1,vLLM 首日就能跑,上下文到 52 万 token,想本地部署智能体编程的可以看看。
IQuest 团队发布 IQuest-Q1,一个面向智能体编程的 320B MoE 模型,每 token 激活 15B 参数,256 个专家中 8 个在线,上下文长度达 524,288。vLLM 提供首日支持,复用其混合 KV cache 协调器,实现 3 层滑动注意力加 1 层全注意力的组合,88 层中仅 25 层随上下文增长缓存。推理侧采用 EAGLE 投机解码与概率化草稿采样,配合递归 MTP 头。
IQuest-Q1 by @IQuest_research has day-0 support in vLLM: a 320B MoE for agentic coding, 15B active per token, 256 experts with 8 live, 524,288 context. 🎉
Built on pieces vLLM already has: 🔹 The hybrid KV cache coordinator, for a 3 sliding to 1 full attention layer mix, so only 25 of the 88 layers grow a cache with the context 🔹 A sinks path in the attention backends, for the sink attention the checkpoint asks for 🔹 EAGLE speculative decoding with probabilistic draft sampling, which is how the recursive MTP head drafts