USAI-Quant:首个建筑环境定量推理基准
USAI-Quant: A Quantitative Reasoning Benchmark for Vision-Language Models in Built Environments
美国团队推出首个建筑环境定量推理基准USAI-Quant,测试了335个城市数据,发现当前VLM在数值推理上普遍不足。
美国研究人员创建了USAI-Quant基准,这是首个用于评估视觉语言模型在建筑环境中定量推理能力的基准。该基准基于美国335个最大城市的高分辨率遥感图像和定量建筑环境指标。研究评估了通用视觉语言模型和遥感视觉语言模型在三个复杂级别上的表现,发现当前最先进的模型在数值推理任务上表现不佳。
USAI-Quant: A Quantitative Reasoning Benchmark for Vision-Language Models in Built Environments
Large vision-language models (VLMs) have emerged as a powerful paradigm for urban and spatial AI. However, current state-of-the-art large VLMs still struggle with quantitative reasoning on remote sensing imagery. Existing benchmarks and algorithms are predominantly based on qualitative Visual Question Answering (VQA), providing limited insights into the quantitative reasoning capabilities of VLMs for built environment metrics. To address this gap, we develop Quantitative Urban and Spatial AI benchmark (USAI-Quant), the first benchmark designed to quantitatively evaluate VLM's reasoning capabilities on built environment metrics via remote sensing imagery. USAI-Quant is curated from the 335 largest U.S. cities, aligning high-resolution remote sensing images with quantitative built environment metrics. We then evaluate both general-purpose and remote sensing VLMs (RS-VLMs) by applying VQAs to tens of built environment metrics across three complexity levels. Our results reveal that current state-of-the-art models consistently fall short on numeric reasoning tasks. We further conduct in-depth analyses across models, question types, and geographic locations, uncovering insights into performance variability and task-specific challenges.