这篇论文揭示了AI如何在没有明确标签的情况下对艺术作品进行美学分类,与人类审美有何不同。
研究人员提出了一种自监督框架,将文本、音频、图像和视频四种模态投影到共享的256维嵌入空间中。该框架通过迭代聚类发现AI形成的美学结构,并分析了AI生成的聚类分配与人类情感标签之间的差异。这项研究有助于理解AI如何构建跨模态相似性,为检索增强生成(RAG)系统组织异构媒体集合提供支持。
How AI Experiences Art: Emergent Aesthetic Structure in a Self-Supervised Multimodal Embedding Space
Aesthetics are an important part of the symbolism of artistic works. Although subjective, humans categorize art based on the emotion evoked regardless of modality. What remains under-explored is how AI models form their own aesthetic categorization of human-produced media without explicit labels or cross-modal supervision. We present a self-supervised framework that projects four modalities (text, audio, image and video) into a shared 256-dimensional embedding space and applies iterative clustering to discover aesthetic structure. We discuss the divergence between AI-generated cluster assignments and human affective register labels on a weakly supervised multimodal dataset. This work has applications in understanding how AI structures cross-modal similarity, organizing heterogeneous media collections for Retrieval-Augmented Generation (RAG), and automated data labeling.