论文

LoopCD提升循环Transformer解码效率

Decoding Looped Transformers Better for (Almost) Free

精选理由

LoopCD让循环Transformer在不增加计算成本的情况下大幅提升解码质量,还能减少计算量。

研究人员提出LoopCD,一种无需训练的对比解码框架,通过对比最终预测与早期循环传递来指导token选择。LoopCD-Logits将Ouro-2.6B-Thinking的AIME 2024准确率从61.88%提升至73.33%,LoopCD-Hidden将Huginn的HumanEval准确率从22.56%提升至31.71%。该方法可将循环次数减半,同时保持或超越全深度基线性能,减少22.5%至48.2%的前向计算量。

原文 · arXiv cs.LG

Decoding Looped Transformers Better for (Almost) Free

Looped Transformers achieve parameter efficiency by repeatedly executing a shared block across recurrent loops. Each loop yields an intermediate representation decodable for the same next token, yet standard decoding discards earlier states. Because earlier loops embody less computation, recurrence inherently supplies aligned weak-and-strong prediction pairs without auxiliary models or external training. We introduce LoopCD, a training-free contrastive decoding framework that guides token selection by contrasting the final prediction with an earlier recurrent pass, operating either in logit space with one extra output pass (LoopCD-Logits) or in hidden-state space with zero output overhead (LoopCD-Hidden). Across four looped Transformer families, LoopCD delivers substantial, consistent gains at full recurrent depth: LoopCD-Logits raises Ouro-2.6B-Thinking's AIME 2024 pass@1 from 61.88% to 73.33%, while LoopCD-Hidden lifts Huginn's HumanEval pass@1 from 22.56% to 31.71%. Crucially, these performance gains enable halving the number of recurrent loops while still matching or exceeding full-depth unguided baselines, reducing forward FLOPs by 22.5% to 48.2%. By transforming intermediate recurrent states into effective guidance signals, LoopCD achieves superior decoding quality while substantially reducing inference compute.

  • Aran Komatsuzaki (论文推介)05:12原文