论文精选

MUSE 教育场景多模态理解基准测试集发布

MUSE: Benchmarking Large Vision-Language Models on Multi-Modal Understanding in Situated Education

精选理由

这个新基准测试集 MUSE 很有用,它专门针对教育场景下的艺术理解,能帮你更准确地评估不同模型的实际能力。

MUSE 是一个专门用于评估大模型在艺术教育场景下多模态理解能力的基准测试集。它包含十二个任务,涵盖视觉感知、语义和情感解读、文化理解以及组合推理,并使用精心挑选的艺术图像,特别关注新加坡和东南亚多元文化背景下的内容。评估结果显示,不同模型在这些能力维度上存在显著差异,尤其是在情感解读和组合推理方面。

原文 · arXiv cs.AI

MUSE: Benchmarking Large Vision-Language Models on Multi-Modal Understanding in Situated Education

Large vision-language models have achieved remarkable progress in multi-modal understanding, yet their capabilities in educational settings remain insufficiently evaluated. In AI-assisted language learning, models must interpret artistic imagery, understand its semantic, affective, and cultural content, and reason about visual context to support meaningful interaction. However, existing benchmarks primarily focus on real-world images or domain-specific educational reasoning, providing limited coverage of artistic educational content. To address this gap, we introduce MUSE, a benchmark for evaluating large vision-language models on artistic image understanding in situated educational applications. MUSE decouples image annotation from question generation, enabling diverse tasks with controllable difficulty while reducing annotation effort. It comprises twelve tasks spanning visual perception, semantic and affective interpretation, culture understanding, and compositional reasoning, together with diverse artistic images deliberately curated to center Singaporean and Southeast Asian multicultural contexts alongside Western art traditions, covering multiple themes and difficulty levels. Evaluation of open-source and proprietary models reveals substantial disparities across capability dimensions, particularly in affective interpretation and compositional reasoning. Our analysis further identifies common failure modes and key challenges for developing trustworthy multi-modal models for education. We hope MUSE will serve as a standardized benchmark for advancing multi-modal understanding in situated educational applications.