SciMIF基准为评估MLLMs在科学领域的指令遵循能力提供了新工具,揭示了不同学科间的性能差异,对科学研究和MLLMs发展具有重要意义。
SciMIF是一个新的基准,旨在评估多模态大型语言模型在遵循复杂科学指令方面的能力。通过分析5个代表性科学学科的22个不同任务,提出包含10个约束组的全面分类法。实验发现,不同科学学科间存在显著性能差异,化学对当前MLLMs更具挑战性。SciMIF填补了评估科学领域多模态指令遵循的空白,为MLLMs在严格科学应用中的未来改进奠定基础。数据与代码将在GitHub上发布。
SciMIF: Understanding Multimodal Instruction Following in Scientific Domains
Understanding instruction-following capabilities in scientific domains is essential for effectively leveraging Multimodal Large Language Models (MLLMs) to advance the development of scientific fields. In this work, we introduce SciMIF, a novel benchmark designed to evaluate the capability of MLLMs in following complex scientific instructions. Specifically, based on an extensive analysis of 22 distinct tasks across 5 representative scientific disciplines, we propose a comprehensive taxonomy comprising 10 constraint groups that captures both general functional requirements and discipline-specific characteristics. Guided by this taxonomy, we develop a high-fidelity instruction injection pipeline to systematically augment existing scientific datasets. We conduct comprehensive experiments on multiple state-of-the-art closed-source and open-source MLLMs. Our findings reveal significant performance disparities across different scientific disciplines, with chemistry posing greater challenges for current MLLMs. Furthermore, we observe that increasing the model scale does not yield corresponding improvements in constraint adherence, and current models still struggle severely with fine-grained constraints and instructions requiring the deep application of disciplinary knowledge. SciMIF fills the current void in evaluating multimodal instruction adherence within scientific domains, laying a crucial foundation for future enhancements of MLLMs in rigorous scientific applications. Data and code will be released at https://github.com/shenye7436/SciMIF .