PathView-Bench:评估多模态大模型病理图像细粒度多尺度理解

PathView-Bench: Can Multimodal Large Language Models Achieve Fine-grained Multiscale Understanding of Pathology Images?

精选理由

PathVU这个新基准测了18个多模态模型,发现它们在病理图上连定位和数数都容易出错,搞医学AI的值得看看。

AI 摘要

PathVU是一个面向病理图像细粒度多尺度视觉理解的基准,基于23个公共病理数据集构建,包含14个VQA任务、61,673张图像和308,070个样本,覆盖28个器官。该基准通过Region FOV和Slide FOV两种视野,评估模型的区域定位、视觉识别、数量估计、空间推理和上下文不足判断能力。对18个通用、医学和病理专用的多模态大语言模型评测显示,现有模型在细粒度视觉任务上存在明显局限。

原文 · arXiv cs.AI

PathView-Bench: Can Multimodal Large Language Models Achieve Fine-grained Multiscale Understanding of Pathology Images?

Multimodal large language models (MLLMs) are increasingly used to analyze pathology images. However, dominant multimodal benchmarks in pathology mainly score final diagnostic answers, captions, or reports. These evaluations provide limited insight into whether a model understands the multiscale visual content needed for pathology reasoning and decision-making. We introduce PathVU, a vision-anchored benchmark for fine-grained and multiscale visual understanding in computational pathology. Built from 23 public pathology imaging datasets with human-supervised labels and spatial annotations, PathVU evaluates MLLM understanding in two fields of view: Region FOV for high-resolution local regions and Slide FOV for macro whole-slide views. By converting raw annotations into deterministic task targets, PathVU enables programmatic scoring of region localization, visual recognition, quantity estimation, spatial reasoning, and insufficient-context judgment. The benchmark contains 14 VQA-style tasks, 61,673 images, and 308,070 samples across 28 organs and 7,253,526 annotations. Evaluating 18 representative general-purpose, medical-domain, and pathology-oriented MLLMs, we observe substantial limitations even in advanced models on fine-grained visual tasks across multiscale pathology images. PathVU provides a reproducible basis for developing and evaluating pathology MLLMs with explicit multiscale visual understanding.