Qwen3.5 达 580 tps 创纪录,TokenSpeed 引擎优化开源 LLM 推理

Fast, faster, Qwen. 🚀 Thrilled to see Qwen3.5 reaching a record-breaking 580 tps for agentic workl...

精选理由

580 tps 意味着智能体应用可以几乎实时响应,做 LLM 推理优化或 Agent 开发的团队值得关注这个开源方案,可以直接参考 PyTorch 博客里的实现细节。

AI 摘要

阿里 Qwen 团队联合多家合作伙伴,在 TokenSpeed 推理引擎上对 Qwen3.5 模型进行极致优化,实现了 580 tokens/秒的推理速度,创下智能体工作负载的新纪录。该成果得益于 NVIDIA GPU、FlashAttention-4 优化以及 PyTorch 社区的支持。这一里程碑展示了开源大模型在推理性能上的巨大潜力,尤其适合对延迟敏感的智能体应用场景。PyTorch 官方博客已发布完整技术细节。

原文 · 阿里通义 Qwen

Fast, faster, Qwen. 🚀 Thrilled to see Qwen3.5 reaching a record-breaking 580 tps for agentic workl...

Fast, faster, Qwen. 🚀 Thrilled to see Qwen3.5 reaching a record-breaking 580 tps for agentic workloads on the TokenSpeed engine! This milestone wouldn't be possible without our incredible partners. Huge thanks to @lightseekorg , @NVIDIAAI , the Mooncake team, and @tri_dao for the pioneering FA4 optimization. Together, we are pushing the boundaries of open-source LLM inference. 🤝✨ Dive into the full @PyTorch blog post below! pytorch.org/blog/up-to-580… cZ #Qwen w #Qwen3_5 3 #TokenSpeed e #LLM L #Inference n #AI # #PyTorch r #OpenSource r #AgenticAI c #HighPerformance nce PyTorch @PyTorch The speed-of-light optimization for Qwen3.5 on the TokenSpeed inference engine is a significant milestone, achieving a record-breaking 580 tokens per second (tps) for agentic workloads on NVIDIA GPUs. In the PyTorch Foundation's latest community blog post, you can learn all about the complete design, implementation, and optimization of Qwen3.5 models in the TokenSpeed inference framework and see for yourself how this work is improving performance 👉 bit.ly/4uGUvIS k This achievement was a joint effort between the @Alibaba_Qwen inference team, @lightseekorg Foundation TokenSpeed team, @NVIDIAAI , and the Mooncake team, with special contributions from @tri_dao for FlashAttention-4 (FA4) optimization. @KVCache_AI 🔗 View Quoted Tweet 💬 23 🔄 33 ❤️ 389 👀 67221 📊 58 ⚡