论文精选

Actor-Critic with Action Chunking (AC2) 方法发布

精选理由

AC2 方法让智能体在未完成轨迹时也能学习,通过动作块分配信用,比传统方法更高效。

研究人员提出 AC2 方法,无需将轨迹滚动至完成即可分配信用。该方法通过学习批评家对每个动作块结束状态进行评分,使策略能够在不观察终端奖励的情况下更新。AC2 使用 10k 个令牌的动作块而非单个令牌分配信用,为批评家提供更有意义的轨迹评估部分。

原文 · Tanishq Abraham (论文推介)

Trust the Critic More

"We introduce Actor-Critic with Action Chunking (AC2) that removes the need to roll every trajectory to completion. AC2 instead assigns credit to action chunks: short continuations of prefixes of past trajectories. A learned critic scores the state reached at the end of each action chunk, allowing the policy to update without observing a terminal reward."

"First, we introduce local readiness which uses critic-based updates on a problem only when the critic is sufficiently accurate on that particular problem. Second, when available, we provide the critic with a reference solution from a previous successful rollout. Third, we assign credit over action chunks of 10k tokens rather than individual tokens, giving the critic a more meaningful portion of the trajectory to evaluate."

Link: https://t.co/UxBYsKAvHG