论文精选73°

PAWBench评估视频生成模型概率对齐能力

PAWBench: How Far Are We from Probabilistically Aligned World Modeling?

精选理由

PAWBench首次系统评估视频生成模型作为世界模型的概率对齐能力,揭示了当前模型与理想状态间的差距。

AI 摘要

PAWBench基准测试评估了11个当前视频生成系统在50个场景下的表现。研究发现,没有模型能够同时匹配参考概率并恢复有效行为范围。该研究引入了概率对齐概念,将世界模型定义为可能行为的分布采样器。PAWEval协议将重复视频展开转换为可能物理行为的经验分布。

原文 · arXiv cs.AI

PAWBench: How Far Are We from Probabilistically Aligned World Modeling?

Recent video generation models are increasingly framed as world models. Many physical processes can unfold in more than one valid way. Therefore, a world model should reproduce not only a plausible trajectory, but also the distribution of possible behaviors under the same initial observation and action. We call this distribution-level requirement probabilistic alignment. However, existing evaluations largely assess individual-video plausibility and do not test whether repeated generations recover the correct distribution. This raises a central question: how far are current video generators from probabilistically aligned world modeling? To answer it, we formalize probabilistic alignment as a distributional criterion for world models and introduce PAWBench, a benchmark for evaluating video generators as stochastic samplers of world dynamics. We further introduce PAWEval, an outcome-level protocol that converts repeated video rollouts into empirical distributions over possible physical behaviors. Across 50 scenarios and eleven current systems, no model consistently matches the reference probabilities while recovering the range of valid behaviors. Having established this gap, we test whether language prompts, initial noise sampling, or model training can reshape the model's predictive distribution. We believe our work can serve as a foundation for future efforts to move towards probabilistically aligned world modeling.