SQAM:用标量伴随匹配做 Q-Learning,加速流策略 RL 微调
Q-Learning with Scalar Adjoint Matching
流策略微调太慢的问题有解了,这篇 SQAM 砍掉了每步的反向传播计算,OGBench 难题成功率提升最多 35 个点,还验证到了真实双臂机器人上。
论文提出 SQAM,针对流策略(flow policy)离线 RL 微调中 adjoint matching 需要每步做向量-雅可比积的高成本问题,利用预训练流策略速度雅可比集中于对角线的观察,推导出闭式标量伴随,按流时间缩放最终动作处的价值梯度。SQAM 在 OGBench 最难的四个域上,成功率比各域最强基线高出 18 到 35 个百分点。团队还在真实双臂机器人上微调视觉-语言-动作策略,SQAM 在全部三个任务上优于监督微调。
Q-Learning with Scalar Adjoint Matching
Flow policies capture rich and diverse action distributions, and fine-tuning them with off-policy RL to improve beyond the demonstrations has drawn growing interest. However, fine-tuning a flow policy against a learned value function is not trivial, because the policy generates its action over many flow steps. Adjoint matching offers a principled way to update the flow model itself by propagating value information from the final action back to each flow step, but it requires a vector--Jacobian product through the policy at every step, a cost that grows with the number of flow steps and the policy size. We observe that the batch-averaged velocity Jacobian of pretrained flow policies concentrates on its diagonal. Motivated by this finding, we derive a closed-form scalar adjoint that scales the value gradient at the final action by the flow time, eliminating the per-step vector--Jacobian products. We further find that controlling the critic's value at policy-generated actions is particularly important under the scalar adjoint. Based on these findings, we propose Q-learning with Scalar Adjoint Matching (SQAM), which combines the scalar adjoint with a value penalty at those actions. SQAM's gains concentrate on the four hardest OGBench domains, where its success rate exceeds that of the strongest baseline in each domain by 18 to 35 percentage points. To test whether SQAM extends to large pretrained policies, we also fine-tune a vision-language-action policy on a real bimanual robot. SQAM improves over supervised fine-tuning on all three tasks.