论文精选

Looped 模型训练新方法:利用不动点实现 KV 共享与加速

Towards Looped Models Done Right, Part II: Rethinking at Fixed Points

精选理由

如果你关心 looped 模型的训练效率,这篇论文给出了一套实在的改进,KV cache 缩小 3 倍还不掉点。

looped 语言模型每多循环一次,训练、解码、prefill 和 RL 的成本都会增加。论文观察到当循环状态接近不动点时路径不再重要,据此可实现截断反向传播、几乎无损的终端 KV 共享解码、prefill 最高 1.79 倍加速的蒸馏学生模型,以及比回放轨迹反传快 2 倍的 RL 梯度计算。方法上论文改进了深度先验和输入注入两个组件:用预测反馈加熵项学习深度先验,并用正交注入移除状态中沿输入方向的分量。从 100M 到 1.6B 参数规模,学习到的先验和正交注入在困惑度上均优于 Huginn 的先验和现有注入方案。在 1.6B 规模下,学习先验配合缩小 3 倍的 KV cache 即可匹配固定深度训练配完整 cache 的下游平均成绩。

原文 · arXiv cs.LG

Towards Looped Models Done Right, Part II: Rethinking at Fixed Points

Every recurrence of a looped language model adds cost in training, decoding, prefill, and reinforcement learning (RL). The closer recurrent states get to fixed points, the less the path to them matters. This enables truncated backpropagation in training; terminal key-value (KV) sharing for decoding with almost no loss in accuracy; a distilled student that prefills up to 1.79x faster; and RL updates that compute gradients from saved rollout states, 2x faster than backpropagating through the replayed trajectory. We therefore improve the two components of training that shape these fixed points: the depth prior and input injection. Fixed-depth training breaks KV sharing, and Huginn's broad depth prior supports sharing but dilutes supervision at the target depth more than sharing requires; we learn the prior from prediction feedback, with an entropy term that keeps it broad. Existing injection schemes let the state's component along the input amplify or cancel the injection; we remove this component with orthogonal injection. From 100M to 1.6B parameters, the learned prior and orthogonal injection lower perplexity at every scale relative to Huginn's prior and existing injection schemes, respectively. At 1.6B, the learned prior with a 3x smaller KV cache matches the downstream average of fixed-depth training with the full cache.