论文

KV²:先粗筛再精评的 KV 缓存压缩方法

KV$^2$: A Self-Refining KV Cache

精选理由

长上下文缓存太吃内存的话看看这篇,KV² 不用重跑整个 prompt,预算压到 2% 还能比基线高 40 多分。

KV² 是一种面向可复用场景的 KV 缓存压缩方法,采用两阶段策略:先用轻量代理评分器挑出有信息量的 token,再只对这一小部分做重建评分。在 RULER 16K 基准、2% 缓存预算下,它比次优基线的平均分高出超过 40 个百分点;在 LongBench 上,2%-10% 预算区间内平均分最高,同时压缩阶段的运行时间和峰值内存低于全上下文重建方案。论文代码已在匿名仓库开源。

原文 · arXiv cs.AI

KV$^2$: A Self-Refining KV Cache

The memory footprint of the key-value (KV) cache constrains the practical use of long-context models, and it dominates cost when one prefilled context must later serve many different queries. In this reusable setting, query-agnostic compression trades cost against quality: lightweight estimators are cheap but less accurate, whereas full-context reconstruction scoring is more accurate yet reprocesses the entire prompt. We introduce KV$^2$, a query-agnostic KV-cache compression method based on selective reconstruction. KV$^2$ first uses a lightweight proxy scorer to identify informative in-context tokens, then reprocesses only this subset to compute final eviction scores. On RULER, Needle-in-a-Haystack, and LongBench, KV$^2$'s margin over baselines widens as the budget tightens: on RULER 16K at a 2% KV-cache budget it improves the average score over the next-best baseline by more than 40 percentage points, and on LongBench it attains the highest average across 2%-10% budgets at lower compression-stage runtime and peak memory than full-context reconstruction. Reusable KV-cache compression thus does not require reprocessing the full context. Our code is available at https://anonymous.4open.science/r/KVsquared-0B97.