论文

DRIVE:用多样性驱动的 RL 微调提升 VLA 泛化能力

Many Ways to Succeed: Diversity-Driven RL Fine-Tuning for VLA Generalization

精选理由

一篇机器人方向的新论文:给 VLA 模型加了个多样性奖励来做 RL 微调,π_0 分布外成绩提了 5.3 分,真机上成功率也从 64% 涨到 73%。

论文提出 DRIVE,将成功行为的多样性显式转化为强化学习目标,用于微调视觉-语言-动作(VLA)策略。方法上,DRIVE 在相同任务条件下对 rollout 分组,通过时间对齐比较轨迹,并基于相对行为差异构建成功条件下的内在奖励。在 LIBERO-Plus、ManiSkill3 和 RoboTwin 2.0 三个基准上,DRIVE 使 π_0 的平均分布外性能比普通 RL 微调提升 5.3 分,π_0.5 提升 2.0 分。在双臂 AgileX PiPER-X 实体平台上,平均分布外成功率从 64.1% 提高到 73.3%,提升 9.2 个百分点。

原文 · arXiv cs.LG

Many Ways to Succeed: Diversity-Driven RL Fine-Tuning for VLA Generalization

Reinforcement learning (RL) fine-tuning improves vision-language-action (VLA) policies through closed-loop experience, yet generalization beyond the fine-tuning distribution remains limited. Our analysis reveals a selective reshaping of exploration: RL contracts behavior globally, yet diversifies successful trajectories, elicits success with fewer rollouts, and covers more of the latent task-valid solution space than supervised fine-tuning. Broader successful-mode coverage may provide alternative strategies under distribution shifts. Inspired by this, we introduce DRIVE (Diversity-driven RL fIne-tuning for VLA gEneralization), which turns successful-behavior diversity into an explicit RL objective. DRIVE groups rollouts under matched task conditions, compares their trajectories with temporal alignment, and derives a success-conditioned intrinsic reward from relative behavioral diversity. This design encourages broader coverage of feasible solutions without rewarding diverse failures or superficial timing differences. Across LIBERO-Plus, ManiSkill3, and RoboTwin 2.0, DRIVE improves the average out-of-domain (OOD) performance over vanilla RL fine-tuning by 5.3 points on $π_0$ and 2.0 points on $π_{0.5}$. On a dual-arm AgileX PiPER-X platform, DRIVE further increases average OOD success from 64.1% to 73.3% (+9.2 points), demonstrating gains that persist under physical deployment.