iS-KV:基于块增量 SVD 的在线 KV 缓存压缩方法
iS-KV: Online Low-Rank KV Cache Compression via Block-Incremental SVD
长链推理时 KV 缓存会爆炸,这篇提出用增量 SVD 压到 4-5 倍还能保住准确率,比删 token 的老办法稳,做推理部署的可以看看。
iS-KV 是一种面向长链推理的在线 KV 缓存低秩压缩方法,保留近期窗口的精确表示,同时将较早状态折叠进有限秩的表示,并在基更新时同步历史坐标以避免漂移。在 DeepSeek-R1-Distill-Llama-8B 上,iS-KV 在 4.06 倍持久 KV 压缩下达到 82.6% 准确率,接近原模型的 83.6%。在 Qwen3-8B 上,它在 5.64 倍压缩下达到 89.2% 准确率。在相同内存预算下,iS-KV 持续优于基于 token 逐出的压缩基线。
iS-KV: Online Low-Rank KV Cache Compression via Block-Incremental SVD
Long chain-of-thought reasoning substantially increases KV-cache memory during autoregressive decoding, as every generated token introduces new key and value states and causes the cache to grow linearly with decoding length. Existing KV-cache compression methods typically control this growth through token eviction, but irreversible deletion can remove historical states that later reasoning may need to revisit. SVD-based low-rank compression provides an alternative by retaining all positions with a more compact representation. However, extending it from a fixed prompt cache to online decoding is non-trivial. Through our investigation, we find that if the basis is updated for new tokens while old tokens keep their coordinates in the old basis, the stored history drifts substantially. Based on this observation, we propose iS-KV, an online low-rank KV-cache compression method for long-horizon reasoning. iS-KV keeps a recent window exact while incrementally folding older states into bounded-rank representations. As the low-rank basis evolves, it synchronizes historical coordinates with the updated basis to maintain representation consistency. On DeepSeek-R1-Distill-Llama-8B, iS-KV achieves 82.6% accuracy at 4.06-fold persistent-KV compression, close to the original model's 83.6%. On Qwen3-8B, it achieves 89.2% accuracy at 5.64-fold compression. Under matched memory budgets, iS-KV consistently outperforms token-eviction baselines.