论文

论文提出 SLIM 基准与逆动力学损失,修复 JEPA 世界模型规划失效

Keeping JEPA World Models Plannable When Little of the Frame Moves

精选理由

JEPA 世界模型在画面几乎不动时规划会崩,这篇论文用一个逆动力学辅助损失把成功率从 0.003 拉到 0.35,还给了个不需要环境访问的探测方法,做世界模型的人应该看看。

arXiv 论文 2610.03137 提出 SLIM 推物基准,在多物体场景上配对视觉目标与语言目标。在 SLIM 上,能解 PushT 的 LeWM 世界模型成功率不足 1%,探针显示编码器潜变量对动作几乎不敏感。加入一个逆动力学辅助损失后,成功率从 0.003 升至 0.35,硬难度推物档达到 0.16,且 PushT 表现在两倍训练时域上提升。在修复后的潜变量上,小型语言目标头达到导航 0.84(视觉目标 oracle 为 1.00),将推物任务拆成阶段句子后,中、难档成功率从 0.04 提升至 0.25,与目标帧 oracle 持平。

原文 · arXiv cs.AI

Keeping JEPA World Models Plannable When Little of the Frame Moves

Specifying a goal in language rather than as a goal frame is a natural interface for planning with a latent world model, but testing it needs scenes in which language must discriminate between several objects. We build SLIM, a pushing benchmark with several small objects and paired visual and language goals on identical scenes. On SLIM a LeWM world model that solves PushT succeeds on under 1% of trials, although a scripted controller with simulator state solves every tier. Probes locate the failure in the encoder: its latent is nearly action-insensitive, neither pusher nor object positions can be decoded from it, and rollouts are no better than copying the current latent forward. One inverse-dynamics auxiliary loss, applied to encoder latents and to predicted latents through a shared head discarded at test time, restores every probe and raises success from 0.003 to 0.35 (0.16 on the hard pushing tier, where a goal-agnostic policy scores zero), and improves PushT at twice the trained horizon. Controls attribute the repair to the gradient into the encoder, and a response sweep shows that the vanilla model plans once enough of the frame responds to actions. A cheap action-sensitivity probe, computable without environment access, acts as an empirical necessary condition: all configurations below its threshold failed to plan. On the repaired latent, a small language-goal head plans from sentences without retraining the world model: it reaches 0.84 on navigation (visual-goal oracle 1.00), follows the named zone when it is swapped with a decoy, and degrades gracefully to unseen nouns. A single goal sentence rarely completes a push, but given the push as a sequence of stage sentences the head raises success on the medium and hard pushing tiers from 0.04 to 0.25, on par with the goal-frame oracle, also when the switch between stages is read from the latent alone.