STEPQuant:Delta规则循环状态量化研究
STEPQuant: When and Where Errors Matter in Delta-Rule Recurrent State Quantization
清华团队提出STEPQuant,根据误差大小和内存寿命分配精度,在Qwen和Kimi模型上显著降低内存使用。
STEPQuant是一种针对Delta规则循环状态的空间-时间后训练量化框架。该研究在Qwen3.8-27B和Kimi-Linear-48B-A3B-Instruct模型上进行了测试,在6位精度下接近FP32精度,4位配置优于均匀INT8。集成到SGLang后,6位STEPQuant实现了超过5倍的循环状态压缩,将服务内存减少高达68.7%。
STEPQuant: When and Where Errors Matter in Delta-Rule Recurrent State Quantization
Linear attention replaces growing KV caches with fixed-size recurrent states, yet these persistent states can become a substantial memory bottleneck under concurrent serving. Directly quantizing recurrent states to low precision often leads to severe accuracy degradation, as quantization errors propagate through successive state updates. We discover that the impact of these errors depends on two complementary dimensions: temporally, errors in long-lived memory can persist across many decoding steps; spatially, errors in different key rows affect model outputs differently, while state magnitudes vary substantially along both rows and columns. Motivated by these observations, we propose STEPQuant, a spatial-temporal post-training quantization framework for Delta-rule recurrent states. STEPQuant allocates precision according to error magnitude and memory lifetime, and jointly fits key-row and value-column scales based on state distributions and key-row impact on output error. Experiments on Qwen3.8-27B and Kimi-Linear-48B-A3B-Instruct across both long- and short-generation benchmarks show that STEPQuant closely matches FP32-state accuracy under a nominal 6-bit budget and outperforms uniform INT8 in its 4-bit configuration. Integrated into SGLang with optimized GPU kernels, 6-bit STEPQuant achieves over 5x recurrent-state compression and reduces total serving memory by up to 68.7%. Our code is available at https://github.com/Dreamer-Toby/STEPQuant.