这个AI框架能把绘画视觉特征转化为描述术语,连接图像数据与人文语义,为艺术史研究提供新工具。
研究人员提出了一种AI框架,可自动分析绘画风格。该系统使用视觉变换器(ViT)在带有元数据的大型绘画语料库上训练,通过稀疏字典学习将特征分解为共享集合。大型语言模型(LLM)检索相关艺术品及其策展人文本,合成反映风格属性的描述。最后,自主协调LLM应用推理-行动(ReAct)框架,将特征整合为连贯的艺术品描述或比较。
Conducting Stylistic Analysis of Paintings through an Art-History Agent
Attributing an artwork to an artist has traditionally relied on detailed visual observations and descriptions, known as stylistic analysis in art history. By contrast, current artificial intelligence (AI) models used in the field offer only unexplained probabilistic classifications. To bridge this methodological gap, we present an AI framework that automates stylistic analysis of paintings, providing a foundation for enhancing evidence collection, discovery, and verification. By training a vision transformer (ViT) on a large corpus of paintings with metadata, our system encodes this art history-specific data as embeddings. These representations are factorized via sparse dictionary learning into a shared set of features that recur across the training set. A large language model (LLM) then interprets each feature by retrieving associated artworks and their accompanying curator-written texts, and synthesizes them into descriptions that reflect their stylistic attributes. Finally, an autonomous coordinator LLM applies a reasoning-and-action (ReAct) framework to weight, test, and refine these features into cohesive descriptions of an artwork, or comparisons of artworks. This approach converts detailed visual features into descriptive terms, addressing a key challenge in art history. It thus connects the use of images as data with the semantic concerns of humanists, establishing vision-based computational art history as an area for future growth.