论文提出 SLIM 基准与逆动力学损失,修复 JEPA 世界模型规划失效
Keeping JEPA World Models Plannable When Little of the Frame Moves
JEPA 世界模型在画面几乎不动时规划会崩,这篇论文用一个逆动力学辅助损失把成功率从 0.003 拉到 0.35,还给了个不需要环境访问的探测方法,做世界模型的人应该看看。
arXiv 论文 2610.03137 提出 SLIM 推物基准,在多物体场景上配对视觉目标与语言目标。在 SLIM 上,能解 PushT 的 LeWM 世界模型成功率不足 1%,探针显示编码器潜变量对动作几乎不敏感。加入一个逆动力学辅助损失后,成功率从 0.003 升至 0.35,硬难度推物档达到 0.16,且 PushT 表现在两倍训练时域上提升。在修复后的潜变量上,小型语言目标头达到导航 0.84(视觉目标 oracle 为 1.00),将推物任务拆成阶段句子后,中、难档成功率从 0.04 提升至 0.25,与目标帧 oracle 持平。
Keeping JEPA World Models Plannable When Little of the Frame Moves
Specifying a goal in language rather than as a goal frame is a natural interface for planning with a latent world model, but testing it needs scenes in which language must discriminate between several objects. We build SLIM, a pushing benchmark with several small objects and paired visual and language goals on identical scenes. On SLIM a LeWM world model that solves PushT succeeds on under 1% of trials, although a scripted controller with simulator state solves every tier. Probes locate the failure in the encoder: its latent is nearly action-insensitive, neither pusher nor object positions can be decoded from it, and rollouts are no better than copying the current latent forward. One inverse-dynamics auxiliary loss, applied to encoder latents and to predicted latents through a shared head discarded at test time, restores every probe and raises success from 0.003 to 0.35 (0.16 on the hard pushing tier, where a goal-agnostic policy scores zero), and improves PushT at twice the trained horizon. Controls attribute the repair to the gradient into the encoder, and a response sweep shows that the vanilla model plans once enough of the frame responds to actions. A cheap action-sensitivity probe, computable without environment access, acts as an empirical necessary condition: all configurations below its threshold failed to plan. On the repaired latent, a small language-goal head plans from sentences without retraining the world model: it reaches 0.84 on navigation (visual-goal oracle 1.00), follows the named zone when it is swapped with a decoy, and degrades gracefully to unseen nouns. A single goal sentence rarely completes a push, but given the push as a sequence of stage sentences the head raises success on the medium and hard pushing tiers from 0.04 to 0.25, on par with the goal-frame oracle, also when the switch between stages is read from the latent alone.