想知道视频生成模型在科学推理上有多靠谱?这个新基准测了 16 个模型,发现视觉好看不等于科学正确,专有和开源差距不小。
Sci-VBench 是一个评估科学领域视频生成的新基准,包含 1,253 个专家标注示例,覆盖 60 个主题和四个学科。该基准要求模型生成具有科学推理和知识基础的视频,而非仅表面视觉合理性。研究显示,非专家评估者和 MLLM-as-Judge 系统与专家判断有较高一致性。对 16 个模型的测试发现,自动感知质量分数相近,但提示接地和科学因果正确性差异显著,且专有模型与开源模型差距明显。这表明视觉真实性的进步尚未转化为对科学和因果动态的可靠建模。
Sci-VBench: Evaluating Knowledge- and Reasoning-Intensive Video Generation in Science Domains
We introduce Sci-VBench, a comprehensive benchmark for evaluating knowledge- and reasoning-intensive video generation across scientific domains. It contains 1,253 expert-annotated examples spanning 60 subjects across four core disciplines: Natural Science, Healthcare, Humanities & Social Sciences, and Engineering. Each example requires models to generate temporally rich videos that demand scientific reasoning and knowledge-grounded synthesis, going beyond surface-level visual plausibility. We further establish a rubric-based evaluation protocol. Our analysis shows that, under this protocol, both non-expert human evaluators and MLLM-as-Judge systems can achieve relatively high agreement with expert judgments, supporting reproducible evaluation at scale. We benchmark 16 frontier proprietary and open-source models and find that, while automatic perceptual-quality scores cluster tightly across systems, performance on Prompt Grounding and Scientific and Causal Correctness varies substantially, with a pronounced proprietary-open-source gap. These findings show that advances in visual realism have not yet translated into reliable modeling of scientific and causal dynamics.