LeapQuant:线性注意力的高效状态量化方法
LeapQuant: Efficient Linear Attention with Accurate Recurrent State Quantization
研究人员提出LeapQuant,解决了线性注意力模型中状态量化的精度损失问题,在保持FP32基线精度的同时显著提升推理速度。
LeapQuant是一种训练免费方法,能在8位递归状态量化下实现接近无损性能。该方法采用每窗口量化技术,通过在窗口末尾一次性量化状态来减少误差累积。同时,它保留状态中的最大异常值作为高精度补偿令牌,并在量化前平滑剩余残差。在Qwen、Kimi和GLM模型家族的实验中,LeapQuant在NVIDIA B200、RTX PRO 6000和RTX 5090 GPU上实现了2.05-3.70倍的内核级加速和1.47倍的端到端推理加速。
LeapQuant: Efficient Linear Attention with Accurate Recurrent State Quantization
Recent LLMs increasingly adopt hybrid designs that replace standard attention with linear attention, such as Gated DeltaNet (GDN) and Kimi Delta Attention (KDA). Although they compress the context into a fixed-size recurrent state and substantially reduce the cost of long-context processing, repeatedly reading and updating that state remains a major inference bottleneck. Quantization offers a natural way to reduce this cost, but can significantly degrade model quality, due to the accumulation of rounding errors and the presence of outlier rows and columns in the state. To address these challenges, we propose LeapQuant, a training-free method that achieves near-lossless performance under 8-bit recurrent-state quantization. First, to mitigate error accumulation, we propose per-window quantization, which leaps over a window of tokens and quantizes the state only once at its end. Within a window, outputs are computed from the fixed low-bit state together with high-precision buffered updates. Second, to reduce the error introduced by each quantization, LeapQuant retains the state's largest outliers as a few high-precision Compensator Tokens, which share the update path of real tokens. We then smooth the remaining residual before quantization to further reduce the error. Comprehensive experiments across the Qwen, Kimi, and GLM model families show that LeapQuant substantially reduces memory and compute costs during inference. With accuracy comparable to the FP32 baseline, it achieves average speedups of 2.05--3.70$\times$ at the kernel level and 1.47$\times$ for end-to-end inference on NVIDIA B200, RTX PRO 6000, and RTX 5090 GPUs.