PermVLA 提出 CAP 方法:用因子分解顺序作正则化训练 VLA 策略
PermVLA: Factorization Order as a Regularizer for VLA Learning
一篇 VLA 训练新论文,思路挺巧:同一批数据换个揭示顺序当正则,LIBERO 和 CALVIN 上都有提升,做机器人策略的可以看看。
论文针对 VLA 策略常用的固定左到右(LTR)因子分解,提出因果锚定置换(CAP)方法,以可调时间前缀采样动作揭示顺序。CAP 通过辅助目标让同一个共享策略从同一专家动作块的不同已知子集预测动作,部署时仍保留确定性 LTR 控制。该方法在不增加演示数据的情况下构造多个条件预测问题,减少对单一时间前缀的依赖。在 LIBERO、LIBERO-Plus 和跨数据集 CALVIN 评测中,CAP 均稳定优于标准 LTR 训练,论文还展示了向扩散动作生成器的扩展。
PermVLA: Factorization Order as a Regularizer for VLA Learning
Vision-language-action (VLA) policies commonly learn action chunks through a fixed left-to-right (LTR) factorization, although the same expert trajectory distribution admits many valid chain-rule factorizations. We identify factorization order as an overlooked regularization choice and introduce causally anchored permutation (CAP), which samples action reveal orders with a tunable chronological prefix. Its auxiliary objective trains one shared policy to predict actions from different known subsets of the same expert chunk, while deployment retains deterministic LTR control. We call this conditional-set augmentation: it creates multiple conditional prediction problems from one expert chunk without adding demonstrations. This discourages reliance on the single chronological prefix used by ordinary teacher forcing. Controlled experiments show that CAP consistently outperforms standard LTR training on LIBERO and LIBERO-Plus, with the same advantage appearing in cross-dataset CALVIN evaluation. A diagnostic that measures the expected squared difference between a chunk's joint log likelihood under two reveal orders verifies that CAP training internalizes agreement across reveal orders. These findings position sampled subset-conditioned auxiliary objectives as a general recipe for constructing VLA regularizers, illustrated by an extension to diffusion action generators.