VBVR-Pro提供了可验证的奖励评分器,使原生视觉推理变得可训练、可验证、可优化和可实验控制,是进行视觉推理研究的强大工具。
VBVR-Pro是一个闭环测试平台,使原生视觉推理通过生成变得可训练、可验证、可优化和可实验控制。该套件包含300个程序生成的任务,模型在VBVR-Pro上训练后,在RISE-Video、MME-CoF-Pro和BabyVision等七个外部视觉推理基准上表现出强大的迁移能力。VBVR-Pro提供可验证的奖励评分器,用于基于任务的评估,并通过系统研究领先的MLLM作为评委,识别了普遍的VLM作为评委范式的重复失败模式。此外,VBVR-Pro还支持对超过30个图像、视频和交错生成器的可控模态研究,分析显示视频生成在需要持续时空状态跟踪的任务中表现最强,而交错生成提供了计算高效的替代方案。
VBVR-Pro: A Scalable and Verifiable Suite for Native Visual Reasoning
Native visual reasoning treats visual generation as the medium of reasoning itself: visual states (i.e. images and videos) are not merely inputs to be understood or outputs to be rendered, but first-class substrates for problem solving beyond language. Yet progress remains bottlenecked by the lack of scalable training tasks, reliable feedback, and controlled comparisons across generative substrates. In this work, we introduce VBVR-Pro, a closed-loop testbed that makes native visual reasoning through generation trainable, verifiable, optimizable, and experimentally controllable. 1) Task scaling. VBVR-Pro turns visual reasoning into a controlled task space of 300 procedurally generated tasks. Models trained on VBVR-Pro show strong transfer beyond the proposed suite across seven external visual reasoning benchmarks such as RISE-Video, MME-CoF-Pro, and BabyVision. 2) Verifiable rewards. VBVR-Pro provides verifiable reward scorers for task-grounded evaluation. Through a systematic study of leading MLLMs as judges, we identify recurring failure modes of the prevalent VLM-as-a-judge paradigm. In contrast, the proposed scorers are grounded in deterministic, task-specific rules, achieve fine-grained alignment with human judgments. Importantly, they serve as reliable reward signals for large-scale multi-task reinforcement learning and demonstrate stronger post-RL performance across visual reasoning tasks. 3) Mechanism study. VBVR-Pro enables controlled modality studies across more than 30 image, video, and interleaved generators. Our analysis shows that video generation remains strongest for tasks requiring persistent spatiotemporal state tracking, while interleaved generation provides a compute-efficient alternative. Critically, ablations and probing suggest the presence of vision-native trajectories that are crucial to visual reasoning. We release all data, models, scorers, and code.