这篇论文没吹牛,诚实地告诉你RMM在标准基准上跟H2O差不多,甚至不如SnapKV,但解释清楚了为什么,适合想搞内存优化的人看看。
本文提出将语言模型有界工作记忆的驱逐问题重新定义为对隐藏信号的估计,引入固定滞后平滑(fixed-lag smoothing)视角,并实例化为无需训练的RMM策略,它是H2O的严格泛化。在控制实验中,RMM能比累积注意力更好地识别有用记忆,但在NVIDIA KVPress基准上,RMM与H2O性能相当,在多轮对话中甚至输给SnapKV和H2O。原因在于自然文本中模型对大多数标记的预测正确,使得演示效用与累积注意力接近。本文贡献在于框架及对何时测量优于积累的诚实分析。
Eviction as Estimation: A Fixed-Lag Smoothing View of Test-Time Memory, and When Measuring Beats Accumulating
A language model with a bounded working memory must repeatedly decide which stored items to keep. Every deployed method decides the moment an item arrives, from the past (StreamingLLM, H2O) or from a guess about the future (SnapKV). We recast the choice as an estimation problem on a hidden signal, whether an item will be reused, placing existing methods on one axis, the commit lag $H$: online filters and learned predictors commit at $H=0$, while Belady's offline optimum sits where the whole future is known. The missing regime in between, fixed-lag smoothing, waits a bounded number of steps, observes which items a correct near-future prediction attended to, and only then commits. This measurement, demonstrated utility, turns Belady's unobservable future request into something we read off the model itself. We instantiate it as a training-free policy, RMM, a strict generalization of H2O that reduces to it exactly when the measurement is uniform. In controlled settings where reuse is endogenous and separated in time, demonstrated utility identifies used memory far better than accumulated attention, and a small bounded memory behaves like a much larger one. But on independent third-party benchmarks, run inside NVIDIA's KVPress harness against its own SnapKV, H2O, and StreamingLLM implementations, the advantage mostly disappears: RMM is on par with H2O for single-turn question answering and loses to both H2O and SnapKV in a streaming multi-turn setting. The cause is simple: on natural text the model is correct about most tokens, so weighting attention by correctness barely changes it, and demonstrated utility collapses onto accumulated attention unless reuse is sharp and endogenous, which standard benchmarks do not exercise. Our contribution is the framework and an honest map of when measuring beats accumulating, not a new state of the art.