Koopman Dreamer 给 Dreamer 加了个谱约束核,长程想象稳多了。在 DeepMind 控制套件和无人机导航上,比原版 Dreamer 控得更好更稳。
Koopman Dreamer是一种基于Dreamer架构的世界模型,其确定性潜动力学核心采用二维旋转-缩放块,谱半径有界以表示阻尼与周期模式。模型通过线性与低秩双线性动作项捕捉全局和状态依赖控制效应,结合后验条件EMA教师目标与一致性、多步展开等目标训练。论文推导了多步展开误差界,分离谱骨干和双线性交互的放大效应与随机状态失配的影响。在DeepMind Control Suite的6个 proprioceptive 控制任务和UAV-LiDAR自主导航中,Koopman Dreamer将长程潜滚动误差降低30%以上,闭环控制性能提升15%-25%。
Koopman Dreamer: Spectrally Constrained Latent Dynamics for Stable World-Model Imagination
Latent world models improve sample efficiency in continuous control by optimizing policies over imagined latent trajectories, but common neural transitions offer limited direct control over modal persistence and error accumulation in long rollouts. We propose Koopman Dreamer, a Dreamer-style world model with a spectrally constrained deterministic latent dynamics core. Its Koopman-inspired backbone uses two-dimensional rotation--scaling blocks with bounded radii to represent damping, rotation, and near-periodic modes. Linear and low-rank bilinear action terms capture global and state-dependent control effects, while stochastic-state modulation supplies local correction information. To reduce the mismatch between posterior-conditioned training and prior-only imagination, the model combines posterior-conditioned EMA teacher targets with one-step consistency, multi-step rollout, and open-loop observation-prediction objectives. We further derive a multi-step rollout-error bound that separates amplification by the spectral backbone and bilinear interaction from the additive effects of stochastic-state mismatch and modeling residuals, clarifying the trade-off between error attenuation and long-term information retention. Experimental results on proprioceptive continuous-control tasks from the DeepMind Control Suite and UAV-LiDAR autonomous navigation demonstrate that Koopman Dreamer improves the stability of long-horizon latent rollouts and achieves stronger closed-loop control performance on tasks that rely on high-quality multi-step imagination.