论文精选

ArtiFact:65万条多模态文化遗产数据集发布

ArtiFact: A Large-Scale Multi-Modal Cultural Heritage Dataset

精选理由

做多模态数据管理、文化遗产数字化或数据质量研究的团队,这个真实世界的大规模基准能帮你测试模型在细粒度错误检测和语义查询上的真实水平,值得跑一跑。

AI 摘要

数据库社区缺乏结合表格、文本和图像的大规模真实数据集。研究者从大都会艺术博物馆、芝加哥艺术博物馆和荷兰国立博物馆收集了651045条博物馆记录,构建了多模态文化遗产数据集ArtiFact。该数据集包含130209条注入七类错误(如材料时代错乱、时间偏移)的记录,用于跨模态错误检测任务。实验表明,当前系统难以检测领域特定的细微错误,且在语义查询处理中,对文化邻近性、模糊对象类型和历史术语的查询表现不佳。ArtiFact为多模态数据管理研究提供了具有挑战性的基准。

原文 · arXiv cs.AI

ArtiFact: A Large-Scale Multi-Modal Cultural Heritage Dataset

Multi-modal data management has emerged as a central research topic in the database community, spanning data integration, semantic query processing, and data quality assessment. Despite this growing interest, the community lacks large-scale, real-world datasets combining tables, text, and images. We present ArtiFact, a multi-modal cultural heritage dataset of 651045 museum records collected from the Metropolitan Museum of Art, the Art Institute of Chicago, and the Rijksmuseum. We demonstrate the utility of ArtiFact through two downstream tasks. For cross-modal error detection, we introduce a curated taxonomy of seven error categories injected into 130209 records and show that reliably detecting subtle domain-specific errors such as material anachronisms and temporal shifts remain an open challenge. For semantic query processing, we show that current systems struggle with queries involving cultural proximity, ambiguous object types, and historically contingent terminology. Our results position ArtiFact as a challenging benchmark for multi-modal data management research.

ArtiFact:65万条多模态文化遗产数据集发布 · AI 热点