LightSeek的TokenSpeed今天支持Qwen3.8了,多节点下比TP16快三成多,跑2.4T大模型延迟更低。
LightSeek的推理引擎TokenSpeed已支持Qwen3.8(2.4T参数)的大规模多节点部署。作为Day-0开源推理引擎伙伴,LightSeek针对NVIDIA Blackwell做了跨节点DP/EP扩展优化。相比TP16,TokenSpeed的吞吐性能提升超过30%。DSpark投机解码和单CUDA图优化进一步降低了延迟。
Excited to see Qwen3.8 running at scale with TokenSpeed! 🚀 Light on latency, big on speed. Kudos to...
Excited to see Qwen3.8 running at scale with TokenSpeed! 🚀 Light on latency, big on speed. Kudos to LightSeek for the fantastic Day-0 support! @lightseekorg LightSeek Foundation @lightseekorg We’re proud to be the Day 0 open-source inference engine partner for @Alibaba_Qwen 3.8. To serve this 2.4T-parameter model across multi-node @NVIDIAAI Blackwell inference, we optimized DP/EP scaling across nodes, delivering 30%+ faster performance than TP16, plus DSpark speculative decoding with single CUDA graph optimization👇 lightseek.org/blog/tokenspee… T 🔗 View Quoted Tweet 💬 2 🔄 3 ❤️ 38 👀 3987 📊 5 ⚡