论文精选

视频模型具备因果可写性,可修正物理错误运动

A Chosen Future Can Still Be Rewritten: Causal Writability in Video Models

精选理由

朋友,看到这篇论文了吗?它发现视频模型其实有修正物理错误的能力,比如让一个本该快速运动的物体慢下来后,还能通过简单方法让它恢复快运动。这挺有意思的。

当视频模型生成物理上不正确的运动时,我们证明其并非未学习正确运动,而是学习了但未使用。在红块慢振、蓝块快振的视频上训练后,测试红块快振时,即使模型生成慢速运动,通过从简单物理变量预测的低维编辑也能恢复正确快速运动。这种能力被称为因果可写性。在固定强度下,我们发现一个清晰的深度边界:编辑在边界前改变视频,边界后则不改变。这个闭合标志着写入的承诺。运动信号仍然存在,更强的下游写入可以恢复物理运动,而过度增益会导致超调。早期因果可写性预测训练后期哪些错误会被修正:可写入的错误在更深的网络层中,而持久错误则不会。

原文 · arXiv cs.LG

A Chosen Future Can Still Be Rewritten: Causal Writability in Video Models

When a video model generates physically incorrect motion, did it fail to learn the correct motion, or did it learn it but fail to use it? We show the latter: the correct motion remains available inside the model and can still be made to control the generated video. We train on videos where red masses oscillate slowly and blue masses oscillate quickly, then test a red mass with fast observed motion. Even when the model generates slow motion in this conflicting case, a low-dimensional edit predicted from simple physical variables restores the correct fast motion. We call this ability causal writability. At fixed strength, we find a sharp depth boundary: the same edit changes the video before the boundary but not after it. This closure marks commitment for that write. The motion signal nevertheless remains, and a stronger downstream write can restore physical motion, while excessive gain overshoots. Early causal writability predicts which errors training later corrects: those errors are writable at more network depths than errors that persist. We reproduce both causal writability and its sharp closure in a pretrained 1.3B video model, supporting generality across model scale and training regime.