多模态空间推理的视觉信用审计

Visual Credit Audit for Multimodal Spatial Reasoning

精选理由

这篇论文告诉你,模型答对空间题不一定是因为看了图——VCA方法能揪出这种虚假正确。结果发现十几到二十几个百分点的正确回答其实没用到图像信息。

AI 摘要

VCA方法通过分离文本和空白对照,评估多模态大模型(MLLMs)在空间基准中是否真的依赖图像证据。在四个MLLMs和两个基准上,12.73-26.25%的正确决策未获得图像信用。相同分割的图像排列使依赖正确性下降21.25-47.80点,所有95%置信区间均高于零。固定像素关系对比和3x3证据源阶乘设计表明,空对照无法识别关系响应;在受控正确但未获信用的决策中,关系反转响应覆盖81.57-100.00%,32.11%改变答案。

原文 · arXiv cs.AI

Visual Credit Audit for Multimodal Spatial Reasoning

Closed yes/no spatial benchmarks can reward a correct answer even when the image adds little support beyond no-image contexts. Under a fixed forced-choice interface, Visual Credit Audit (VCA) separates two estimands: whether the benchmark image gives the model's declared decision more support than text-only and blank controls, and whether the model responds to relation-specific visual evidence. The first audit is training- and label-free and does not require an answer flip. Applying labels yields dependence-credited correctness (D-CC); on correct items, it equals same-control gold-aligned positive gain, while prediction alignment extends the audit to errors. Across four open MLLMs and two spatial benchmarks, 12.73-26.25% of decisions are correct yet uncredited. Matched same-split image permutation reduces D-CC by 21.25-47.80 points, with every paired 95% interval above zero. Fixed-pixel relation contrasts and a 3x3 evidence-source factorial show why null controls cannot identify relation response. Among controlled correct-but-uncredited agreement decisions, response to relation reversal spans 81.57-100.00%, while 32.11% pooled change answer. Independently audited outcomes on 108 geometry-compatible edits provide a bounded natural-image correspondence check. VCA thereby decomposes benchmark success into correctness, additional image support, and relation-consistent response.