自动驾驶选择性神经符号推理框架
From Scene Graphs to Answers: Selective Neuro-Symbolic Reasoning for Autonomous Driving
自动驾驶领域新框架,用符号处理+LLM组合提升问答准确率,计数问题效果尤其明显。
研究人员提出了一种查询自适应神经符号推理框架,用于自动驾驶问答系统。该框架基于时空场景图(STSG),将计算分配给符号执行器或大语言模型。在5,916个NuScenes-QA问题测试中,使用GPT-5.4-mini的系统准确率达到80.63%,比纯LLM配置提高5.48个百分点;使用DeepSeek-V4-Flash提高6.64个百分点。计数问题提升最为显著,分别提高10.20和12.61个百分点。
From Scene Graphs to Answers: Selective Neuro-Symbolic Reasoning for Autonomous Driving
Autonomous-driving question answering requires reasoning over structured scene information, yet existing vision-language approaches largely delegate heterogeneous reasoning operations to a single neural inference process. We argue that this uniform strategy overlooks a fundamental distinction: some queries admit exact symbolic solutions, while others require semantic interpretation. We introduce a query-adaptive neuro-symbolic reasoning framework that explicitly allocates computation according to the nature of the query. At its core is a hierarchical Spatiotemporal Scene Graph (STSG) that separates persistent object identities from frame-specific states and represents spatial relations and temporal transitions as explicit directed structures. Given a query, a symbolic executor first attempts to resolve it through exact graph operations; only when symbolic execution abstains is an LLM invoked for semantic reasoning. For these unresolved queries, query-conditioned graph retrieval and evidence filtering preserve relation direction, temporal locality, and object semantics, providing the LLM with compact and verified task-relevant evidence. This design shifts the role of the LLM from a universal reasoning engine to a targeted semantic reasoner, while allowing deterministic computation to be handled exactly and efficiently. We evaluate the framework on 5,916 NuScenes-QA questions across all ten scenes of nuScenes v1.0-mini under an oracle-perception setting. The complete system achieves 80.63 percent overall accuracy with GPT-5.4-mini, improving over the corresponding LLM-only configuration by 5.48 percentage points; with DeepSeek-V4-Flash, the improvement reaches 6.64 points. The largest gains occur on counting questions, with improvements of 10.20 and 12.61 points, respectively. These results show that selective reasoning improves both accuracy and inference efficiency.