论文

VIGOR:基于潜空间一致性实现 MBRL 零样本视觉泛化

VIGOR: Zero-Shot Visual Generalization via Latent-Space Consistency in Model-Based Reinforcement Learning

精选理由

强化学习在换背景、变光照时就崩的问题,这篇用潜空间一致性给出解法,Robosuite 上比第二名的基线高了 43.6%。

VIGOR 是一个面向模型-based 强化学习(MBRL)的视觉泛化框架,针对视觉干扰下潜变量推演误差逐级放大的问题,结合弱到强非对称增强、动力学层一致性和编码器层稳定三个组件。在 DeepMind Control Suite 上,VIGOR 超过第二好的基线 3.4%;在 Robosuite 上领先幅度达 43.6%。消融实验显示,替换默认增强方式后泛化性能依旧保持,说明鲁棒性来自潜空间一致性而非特定增强手段。

原文 · arXiv: Google DeepMind

VIGOR: Zero-Shot Visual Generalization via Latent-Space Consistency in Model-Based Reinforcement Learning

Model-based reinforcement learning (MBRL) achieves strong sample efficiency by planning within learned latent dynamics, yet its performance degrades substantially under unseen visual distractions such as background variations, lighting changes, or camera shifts. Unlike model-free RL, where encoder perturbations affect only single-step predictions, MBRL suffers from a two-level vulnerability: visual distractions first push encoder outputs out of distribution, and these errors then compound through recursive latent rollouts over the planning horizon. We propose visual generalization via latent-space consistency in model-based RL (VIGOR), a framework that enables zero-shot generalization to unseen visual distractions while retaining the sample efficiency of its MBRL backbone. VIGOR integrates three interdependent components: (i) asymmetric weak-to-strong augmentation, which pairs weak-only and weak-to-strong latent views within a single batch; (ii) dynamics-level consistency, which enforces augmentation-invariant transition predictions through direct latent regression; and (iii) encoder-level stabilization, which prevents encoder drift under the cross-augmentation supervision imposed by dynamics-level consistency. Evaluations on the DeepMind Control Suite (DMC) and Robosuite show that VIGOR outperforms state-of-the-art model-free and model-based baselines, surpassing the second-best baseline by 3.4% on DMC and 43.6% on Robosuite. Ablations further show that VIGOR's robustness is augmentation-agnostic: replacing the default augmentation with alternatives from distinct perturbation families preserves strong generalization, confirming that latent-space consistency, not the augmentation choice, drives robustness.