论文

WOVEN:视觉世界建模融入多模态大模型

WOVEN: Weaving Visual World Modeling into Multimodal LLMs

精选理由

WOVEN解决了多模态大模型在空间、物理和时间推理上的缺陷,训练数据可替代30-50%任务数据,效果显著。

研究团队提出WOVEN,一个针对视觉转换推理的训练源和基准,包含36,076个示例,覆盖20种场景、5种动作和8种推理类型。评估38个前沿多模态大模型(如GPT-5.4和Qwen3-VL-235B-A22B)发现,即使最强模型在视觉转换推理上也远低于人类水平。在WOVEN上训练的多模态大模型展现出可迁移能力,仅约2,000个样本的训练集就能在26个外部基准中的22个上提升高达27.3个百分点。

原文 · arXiv cs.LG

WOVEN: Weaving Visual World Modeling into Multimodal LLMs

Multimodal large language models (MLLMs) struggle with spatial, embodied, physical, and temporal reasoning. We hypothesize that these failures reflect a shared deficit in visual transition reasoning, and test whether this capability can serve as a shared training primitive, one that different models can learn from different supervision sources and reuse across different tasks, with a systematic training recipe. Existing benchmarks document these deficits separately but do not support controlled comparisons across scenes, actions, and reasoning operations. We therefore introduce WOVEN, a training source and benchmark for visual transition reasoning that organizes transition supervision by scene, action, and reasoning type, using diverse, realistic rollouts from video-pretrained generative models: 36,076 examples across 20 scene types, 5 action types, and 8 reasoning types. We first evaluate 38 frontier MLLMs (e.g., GPT-5.4 and Qwen3-VL-235B-A22B) and find a substantial and systematic deficit: even the strongest models fall far below humans, and the failures recur across model families and persist with scale. We then train MLLMs at multiple scales on WOVEN and find that they learn a shared capability that transfers broadly: training subsets of only about 2,000 items each collectively improve 22 of 26 external benchmarks by up to 27.3 percentage points, and WOVEN data can replace 30-50% of a task's own training data with comparable accuracy. Controlled comparisons further yield a training recipe for visual world modeling, validated prospectively on held-out benchmarks: select supervision by the reasoning operation it teaches rather than by the actions, scenes, or domains it shows, and prefer larger changes to the visual state for robustness. Our work establishes visual transition reasoning as a reusable foundation for systematic visual world-model training in MLLMs.