CARVE：内容感知高效线性注意力，修正记忆盲门控缺陷

精选理由

这篇论文用简单思路修了GDN-2的三个bug，实测1.3B模型困惑度降了0.18，还省内存和参数，想搞高效注意力的话值得看。

AI 摘要

CARVE提出仅擦除关键轴的注意力机制，解决了GDN-2的三个耦合缺陷：记忆盲门控、值轴擦除掩码浪费参数、无法使用WY形式三角形分块求解器。在1.3B参数、100B token训练下，CARVE在WikiText上达到困惑度15.72（比GDN-2低0.18，4.5-sigma效应）。它在9个常识推理基准上领先所有循环基线，并在RULER检索探针上取得SOTA。该方案仅带来0.4%吞吐开销、13%更低峰值内存和19%更少参数。论文还包含六个形式化定理，涵盖记忆容量、Lyapunov稳定性等。

AI 翻译 · 中文

arXiv cs.AIRecurrent models must forget in order to remember, yet the state of the art decides what to erase without consulting what is stored -- the gate sees only the arriving token, not the memory it is about to modify. This mem…

阅读原文