论文精选73°

多图像物体幻觉基准MIOH发布

Fine-Grained Multi Image Object Hallucination Benchmark

精选理由

MIOH基准帮你看清多模态模型在复杂视觉推理中的真实表现,GPT-5和Gemini-2.5-Pro也不例外。

AI 摘要

研究团队推出MIOH基准,系统评估多图像场景下的物体幻觉问题。该基准包含4项基础任务(存在性、计数、属性、位置),3种多图像推理模式(综合、比较、选择性)和3种对抗压力。测试显示GPT-5和Gemini-2.5-Pro等29个模型均存在不同失败模式。幻觉主要源于多图像间物体表示的整合阶段限制,而非感知失败。

原文 · arXiv cs.AI

Fine-Grained Multi Image Object Hallucination Benchmark

Multimodal Large Language Models (MLLMs) are increasingly deployed in multi-image scenarios requiring complex reasoning across visual contexts. However, current MLLMs remain fundamentally limited by object hallucination-generating plausible yet factually inconsistent descriptions about objects. Existing benchmarks, designed primarily for single-image settings or providing only high-level multi-image assessments, cannot systematically diagnose how visual complexity and reasoning demands trigger hallucination. To address this gap, we introduce MIOH, a fine-grained multi-image object hallucination benchmark that systematically evaluates object hallucination across four foundational tasks (existence, counting, attribute, position) through three multi-image reasoning patterns (comprehensive, comparative, selective) under three controlled adversarial pressures (visual context scale, perceptual difficulty, contextual bias). Through evaluation of 29 models, we reveal that even state-of-the-art systems like GPT-5 and Gemini-2.5-Pro exhibit distinct failure patterns across different reasoning patterns and tasks. Our evaluation reveals that hallucination stems not merely from perceptual failures but from integration-stage limitations when maintaining object representations across multiple images. MIOH provides a controlled framework for analyzing multi-image object hallucination and serves as a critical evaluation tool for developing more reliable multimodal AI systems.