VIALS:生命科学中视觉解释基准

VIALS: A Benchmark for Visual Interpretation of Artifacts in the Life Sciences

精选理由

VIALS基准为视觉语言模型在生命科学领域的应用提供了新的评估标准,有助于识别模型在特定领域的局限性。

AI 摘要

VIALS基准包含161个视觉解释任务,涵盖生物技术行业中实验工作流程中检查的各类视觉艺术品。尽管前沿视觉语言模型能够流畅描述自然图像,但它们无法准确解释这些科学图像,反映出在领域知识和特定领域视觉推理能力方面的局限性。相比之下,具有相关领域专业知识的科学家发现这些视觉解释任务很简单。无法类似解释这些图像的AI在专业生命科学工作流程中将具有有限的应用价值,因为这些艺术品是科学家推理、沟通和做决策的核心。

原文 · arXiv cs.AI

VIALS: A Benchmark for Visual Interpretation of Artifacts in the Life Sciences

In professional life sciences workflows, scientists routinely interpret visual artifacts (gel blots, microscopy images, plasmid maps, flow cytometry plots, molecular structures, ...) to inform research decisions. We introduce VIALS, a visual question-answering benchmark with 161 such interpretation tasks, spanning the types of artifacts examined throughout experimental workflows in the biotech industry (rather than polished figures from publications and textbooks). While frontier vision-language models can now fluently describe natural images, we find that they are unable to accurately interpret these scientific images, reflecting limitations in domain knowledge and domain-specific visual reasoning capabilities. In contrast, scientists with relevant domain expertise find these visual interpretation tasks straightforward. AI that cannot similarly interpret such images will have limited utility in professional life sciences workflows, where such artifacts are central to how scientists reason, communicate, and make decisions.