AUTOPILOT-VQA:面向事故场景的行车记录仪视频问答基准

AUTOPILOT VQA: Benchmarking Vision-Language Models for Incident-Centric Dashcam Understanding

精选理由

CVPR 2026 竞赛推出了 AUTOPILOT-VQA 基准,专门测试 AI 模型对行车记录仪事故视频的推理能力,比普通视觉问答更难更贴近安全。

AI 摘要

AUTOPILOT-VQA 是一个针对行车记录仪视频的视觉问答基准,用于评估视觉语言模型在安全关键事故中的推理能力。该基准覆盖天气光照、交通环境、道路布局、路面状态、标志、涉及实体、事故是否发生、撞击位置和事故可避免性等9个类别。数据集基于真实驾驶事故与近事故场景设计结构化问题,要求模型回答场景属性与事件细节。该基准作为 CVPR 2026 竞赛的一部分发布,旨在推动更可解释、鲁棒和安全敏感的视觉语言系统。

原文 · arXiv cs.AI

AUTOPILOT VQA: Benchmarking Vision-Language Models for Incident-Centric Dashcam Understanding

Recent advances in Vision-Language Models, Large Language Models, and Multimodal Large Language Models have improved autonomous driving tasks such as scene understanding, decision making, trajectory prediction, and visual question answering. However, evaluating whether these models can reliably reason about safety-critical incidents remains challenging. To address this gap, we present AUTOPILOT-VQA, an incident-centric visual question answering benchmark for dashcam video understanding. The dataset evaluates different systems through structured questions designed around real-world driving incidents and near-incidents. The benchmark covers diverse safety-relevant categories, including weather and lighting conditions, traffic environment, road layout, road surface state, signage, involved entities, accident occurrence, impact location, and avoidability-related reasoning. By requiring models to answer grounded questions about both contextual scene properties and event-level incident details, AUTOPILOT-VQA moves beyond object recognition toward temporally grounded, safety-aware reasoning. The dataset is released as part of the AUTOPILOT CVPR 2026 competition and provides a standardized benchmark for assessing the reliability of autonomous driving systems in different scenarios. Our benchmark support developments for more interpretable, robust, and safety-conscious vision-language systems for real-world autonomous driving.