论文

AdaStep:用自适应步级权重改进智能体强化学习训练

AdaStep: Adaptive Step Credit Weighting for Agentic Reinforcement Learning

精选理由

训练长程智能体的朋友可以看看这篇,它解决了步级信用分配噪声大的问题,而且不用加 critic,成本很低,在 ALFWorld 等三个环境都有提升。

论文提出 AdaStep,一种面向长程 LLM 智能体强化学习的自适应步级信用分配方法。它将步级权重计算转化为对潜在步级优势的均方误差估计问题,并推导出基于信噪比的逐状态收缩系数。当回报变化来自所选动作时保留局部信用,当变化主要来自下游随机性时则予以抑制。该方法只需轻量标量计算,无需 critic、额外 rollout 或额外模型推理。在 ALFWorld、WebShop、ScienceWorld 三个环境上配合三种模型骨干测试,均取得稳定超越基线的效果。

原文 · arXiv cs.LG

AdaStep: Adaptive Step Credit Weighting for Agentic Reinforcement Learning

Long-horizon LLM agents are typically trained with sparse outcome rewards, making trajectory-level objectives too coarse to distinguish the contribution of individual decisions. Step-level credit assignment provides finer-grained supervision, but its estimates can be unreliable because observed returns also depend on subsequent actions, environment transitions, and trajectory length. We propose AdaStep, an Adaptive Step-credit weighting method that controls how strongly each group-derived local advantage modifies the trajectory-level signal. We formulate this weighting as a mean-squared-error estimation problem for the latent step advantage and, under an explicit conditional sampling assumption, derive an optimal per-state shrinkage coefficient. The coefficient admits a signal-to-total-variance interpretation: it preserves local credit when return variation is attributable to the selected action and suppresses it when variation is dominated by downstream randomness. AdaStep requires only lightweight scalar computation, with no critic, additional rollouts, or extra model inference. Experiments with three model backbones on ALFWorld, WebShop, and ScienceWorld show consistent improvements over baselines at low computational cost.