论文精选

视觉语言模型为何会遗漏小物体

Why VLMs Miss Small Objects, and When Zooming In Is Safe

精选理由

MIT团队揭秘VLMs遗漏小物体的数学原理,提出图像分解策略,在建筑图纸测试中提升召回率。

研究表明视觉语言模型(VLMs)在处理大图像时经常遗漏小物体。研究团队基于图像界面上的两个关键量S(物体一侧的视觉token数量)和L(调用必须覆盖的内容)建立了理论模型。在177,000次请求和797张图像的测试中,该理论预测了8个VLMs的40种排序,其中61种排序在自然和合成图像上显著成立。在建筑图纸测试中,分解方法将召回率提高了最高0.28。

原文 · arXiv: OpenAI

Why VLMs Miss Small Objects, and When Zooming In Is Safe

Vision-language models (VLMs) often miss small objects in large images. We ask three questions: what limits them, which of these limits better models can remove, and whether the classical way of handling large images, local decomposition, still has a future. We answer them with a theory built on two quantities of the image interface: S, the number of visual tokens across an object's side, and L, the content a call must cover. The limits: a W x H image sent whole within N tokens gives an object of side m at most m*sqrt(N/(WH)) tokens per side. Doubling the token budget N therefore raises the largest S a whole image can reach by only 41%, and seeing an image at the S an object needs costs at least order S^2 tokens whatever the model. If recognition improves gradually with S, any search strategy, zoom agents included, obeys a recall-cost frontier. What better models can change: the S an object needs and how much one call can carry, measured in bits per object found; the coverage cost remains. Decomposition: yes. Assuming only that more tokens per object and less content per call do not hurt on average, splitting an image cannot lower recall if no view zooms out relative to the whole image and views overlap by one object. Neither part of this condition can be dropped; we bound the cost of every such decomposition, and a simple rule approaches the bound as the image grows. We test the theory in about 177,000 requests on 797 images. On controlled images, none of 40 orderings predicted for 8 VLMs is violated. Checked after the fact on drawings, floor plans, natural and synthetic images, 61 of 89 implied orderings hold significantly and 4 fail, all for OpenAI models given more pixels than their default path. On construction drawings the rule never significantly lowered recall relative to the whole image and raised it by up to 0.28. Code and data: https://github.com/shijunzhe/vlm-small-objects