基于感知优势估计的多模态推理强化方法
Reinforcing Multimodal Reasoning via Token-Level Perception-Grounded Advantage Estimation
TPAE方法通过细粒度监督提升多模态模型推理能力,在七个基准测试中表现优异。
研究人员提出TPAE方法,通过视觉依赖性和预测熵两个token级指标分析多模态推理链。该方法在七个基准测试中表现优于现有基线,能提供更稳定高效的多模态推理优化。TPAE利用token级优势估计来调制序列级优势,可集成到多种RLVR框架中。代码已在GitHub开源。
Reinforcing Multimodal Reasoning via Token-Level Perception-Grounded Advantage Estimation
Reinforcement Learning with Verifiable Rewards (RLVR) has improved the reasoning capabilities of Multimodal Large Language Models (MLLMs), yet existing frameworks rely on coarse, sequence-level reward signals that lack the fine-grained supervision over the visually-grounded steps within a multimodal reasoning chain. We investigate this gap through the lens of two token-level metrics: visual dependency (i.e. how much a token's prediction relies on the input image features) and predictive entropy. Our empirical analysis reveals two key findings: (1) correct reasoning chains exhibit a markedly sharper entropy reduction as visual grounding intensifies, compared to incorrect ones; (2) pivotal tokens, those whose misprediction triggers reasoning collapse, are statistical outliers in the joint distribution of visual dependency and predictive entropy derived from correct chains. Motivated by these findings, we propose token-level perception-grounded advantage estimation (TPAE), which estimates token-level advantages by measuring each token's statistical consistency with the vision-entropy patterns of correct rollouts. TPAE leverages this granular score to modulate the sequence-level advantage, producing a fine-grained supervision signal that can be integrated into various RLVR frameworks. Extensive experiments on seven benchmarks show that TPAE consistently outperforms leading strong baselines, yielding more stable and efficient optimization for multimodal reasoning. The code is publicly available at https://github.com/Zhihan72/TPAE.