技巧精选

推理工程大师课:量化、投机解码与GLM-5.2自优化

The Inference Engineering Masterclass: 10x faster models, quantization, speculative decoding, Rubin,...

精选理由

Baseten工程师讲推理优化,量化误差能抵消,GLM-5.2自己优化自己,还有20-200%提速空间。

AI 摘要

在latent.space的推理工程大师课中,Baseten的Philip Kiely与@waterloo_intern讲解了模型训练后如何变成快速可靠的产品。他们指出量化误差可以相互抵消,从而解锁更高吞吐量。推理团队目前仍能挖掘20%–200%的性能增益。视频生成面临二次注意力扩展的瓶颈。GLM-5.2还使用自身重写并优化了服务自己的GPU内核。

图片来源 · Latent.Space
原文 · Latent.Space

The Inference Engineering Masterclass: 10x faster models, quantization, speculative decoding, Rubin,...

The Inference Engineering Masterclass: 10x faster models, quantization, speculative decoding, Rubin, & self-optimizing AI latent.space/p/inference-eng @Baseten @philipkiely and @waterloo_intern explain what actually happens after a model is trained, why turning weights into a fast and reliable product creates an entirely new optimization problem, how quantization errors can cancel out to unlock more throughput, why inference teams are still finding 20–200% performance gains, how video generation runs into a quadratic attention wall, and how GLM-5.2 helped rewrite and optimize the GPU kernels serving GLM-5.2 itself. Your browser does not support the video tag. 🔗 View on Twitter 💬 0 🔄 1 ❤️ 7 👀 897 📊 2 ⚡

推理工程大师课:量化、投机解码与GLM-5.2自优化 · AI 热点