论文

TokenProbe 框架:通过压缩冗余 token 将推理开销降低 76%

On the Token Value Inequality in Efficient Reasoning

精选理由

这篇论文发现推理链里大部分 token 其实是废话,压掉 76% 还不掉分,预算紧张的话值得看看他们的做法。

一篇 arXiv 论文研究 CoT 推理中 token 价值的不均匀性,发现用归一化 log probability 可以区分承载关键推理的核心 token 与低置信度的冗余 token。基于这一发现,作者提出 TokenProbe 框架,并结合 GRPO 目标函数选择性地压缩冗余 token。实验显示该方法在保持推理质量的前提下,将 token 用量比基线减少 76%。在相同推理长度预算下,其表现甚至超过 Gemini-3.1-Pro。

原文 · arXiv cs.AI

On the Token Value Inequality in Efficient Reasoning

Chain-of-Thought reasoning has enabled large language models to achieve substantial performance gains on complex tasks. However, these gains come at the cost of dramatically increased token consumption. This raises a fundamental question: is every token in the reasoning trace equally valuable? We present a diagnostic and optimization framework grounded in a key empirical finding: the value of tokens within a CoT reasoning sequence is highly non-uniform, and this non-uniformity can be effectively characterized by token-level log probability signals. We show that normalized log probability helps distinguish core tokens, which carry structural and decisive reasoning content, from redundant tokens, which are exploratory, low-confidence filler that contributes less directly to the final answer. Building on these findings, we formulate the TokenProbe framework around two empirical findings and one claim: findings identify token value inequality first and then establish TokenProbe as a core-token proxy, and the claim introduces an efficient GRPO objective positing that selectively compressing redundant tokens can yield Pareto improvements in the accuracy-token efficiency space. Empirically, our method preserves reasoning quality while reducing the token usage by 76% of the baseline. Under matched reasoning-length budgets, we show that it can even outperform strong flagship baselines like Gemini-3.1-Pro. Homepage: https://runjia.tech/tokenprobe/.