论文

T-Router:基于参数高效强化学习的推理路由方法

T-Router: Learning Thalamic Routing for Reasoning with Parameter-Efficient Reinforcement Learning

精选理由

T-Router用不到0.5%的参数就让大模型推理能力提升近10个点,比LoRA和全参数训练都强。

研究人员提出T-Router模型,在89.5亿参数骨干模型上仅使用4173万参数(0.466%),通过GSM8K强化学习后达到83.64±1.16的MathAvg分数。相比全参数GRPO的73.79±1.83和LoRA的77.28±1.95,T-Router在相同参数预算下表现更优,AIME准确率从48.33提升至60.56。该模型通过可寻址块变化和递归上下文实现受控计算重用。

原文 · arXiv cs.AI

T-Router: Learning Thalamic Routing for Reasoning with Parameter-Efficient Reinforcement Learning

Parameter-efficient reinforcement learning aims to improve reasoning with a compact trainable interface to a pretrained model. We introduce the Thalamic Router (T-Router), which concentrates adaptation on the reuse of completed computations. A compressed, addressable bank preserves block changes; a depth-recurrent controller conditions their selection and relative-scale writeback. This coupling gives thalamic context-dependent routing a concrete computational form: learn which earlier contributions a receiving layer uses, and with what influence. Correctness rewards train the interface while preserving backbone parameters and layer order. On an 8.95B-parameter backbone, T-Router allocates 41.73M parameters (0.466% of the backbone) and achieves 83.64 +/- 1.16 MathAvg after GSM8K RL, compared with 73.79 +/- 1.83 for full-parameter GRPO across three evaluation rounds. At a comparable parameter budget and with matched retries, it exceeds LoRA's 77.28 +/- 1.95 MathAvg, improving all three task families and raising mean AIME accuracy from 48.33 to 60.56. Capacity-controlled comparisons favor addressable block changes and recurrent context; separate search training extends the interface to tool-mediated reasoning. These results establish controlled computation reuse as an effective route to parameter-efficient reasoning reinforcement learning.