纠正学习到的物理不变量改善世界模型运行

Correcting a learned physical invariant improves world-model rollouts

精选理由

这篇论文揭示了世界模型在处理物理约束时的局限性,为改进模型提供了有价值的见解。

AI 摘要

研究测试了仅通过摆动视频训练的冻结DreamerV3模型是否学会了一个标量,该标量在自身的潜在转换中被视为近似守恒。无标签搜索在独立训练的保守模型中恢复了相同的能量类不变量,而在匹配阻尼模型中找不到类似的变量。在自主运行期间,该量值发生漂移。将潜在状态投影回其初始水平集可以减少所有三个保守模型的运行错误,而匹配随机约束通常会增加错误。这些结果区分了动态意义上的不变量与仅可解码的相关性,并揭示了一个具体的失败模式:世界模型可以从像素中学习物理约束,但在想象未来时违反该约束。

原文 · arXiv cs.AI

Correcting a learned physical invariant improves world-model rollouts

World models can predict video without learning dynamics that they reliably preserve. We test whether a frozen DreamerV3 trained only on pendulum video learns a scalar that its own latent transition treats as approximately conserved. A label-free search recovers the same energy-like invariant across independently trained conservative models, while the same procedure finds no comparable invariant in matched damped models. During autonomous rollouts, this quantity drifts. Projecting the latent state back toward its initial level set reduces rollout error in all three conservative models, whereas matched random constraints usually increase it. These results distinguish a dynamically meaningful invariant from a merely decodable correlate and reveal a concrete failure mode: a world model can learn a physical constraint from pixels yet violate that constraint when it imagines forward.