Video-DR 让 AI 边看视频边上网核实信息,35B 参数干翻 Claude-4.5-Sonnet,还开源了,值得试试。
论文提出 Video-DeepResearch,将多模态智能体从静态图像扩展到连续视频流,要求密集时空定位与开放网页探索结合。在 Video-DR-Bench 基准上,Video-DeepResearch-35B-A3B 平均准确率达 64.0%,超过 Claude-4.5-Sonnet(59.0%)5.0 个百分点,也优于 GPT-5(52.5%)和 Gemini 2.5 Pro(57.5%)。框架采用解耦的感知-探索流程,通过分阶段工具解锁强制跨帧视觉定位后再进行网页检索。训练结合监督微调与 GRPO,30B-A3B 变体达到 59.3%,与 Claude-4.5-Sonnet 相当。代码已开源。
Video-DeepResearch: Towards the Next-Generation Multimodal Deepresearch Agent
We introduce Video-DeepResearch (Video-DR), extending multimodal agents from static images to continuous video streams, a setting that demands dense spatiotemporal grounding coupled with open-web exploration. Preliminary evaluations reveal two critical bottlenecks in current models: (1) modality bias, where agents bypass visual tools in favor of textual search, and (2) parametric knowledge leakage, where models rely on internal memory rather than genuine tool-augmented execution. To address these challenges, we propose Video-DR, featuring a decoupled perception-exploration pipeline with stage-wise tool unlocking that compels exhaustive cross-frame visual grounding prior to web retrieval. Our framework adopts a two-stage training recipe: supervised fine-tuning followed by Group Relative Policy Optimization (GRPO), enabling autonomous exploration that breaks the imitation-learning ceiling. Furthermore, we curate Video-DR-Bench, a human-AI collaborative benchmark comprising 200 complex, multi-hop VQA instances. Empirical results demonstrate that our Video-DeepResearch-35B-A3B establishes a new state-of-the-art of 64.0% average accuracy, surpassing proprietary Claude-4.5-Sonnet (59.0%) by 5.0 points and significantly outperforming GPT-5 (52.5%) and Gemini 2.5 Pro (57.5%). The 30B-A3B variant achieves 59.3%, competitive with Claude-4.5-Sonnet and demonstrating the effectiveness of our training paradigm even at compact scale. Code: https://github.com/Osilly/Vision-DeepResearch.