人类识别AI生成图像中的缺陷

When Composition Doesn't Add Up: Humans Identifying Defects in AI-Generated Images

精选理由

这项研究揭示了AI生成图像中的构图缺陷,并提供了CO-AID数据集,对于优化AI图像生成和提升模型性能有重要意义。

AI 摘要

本研究探讨了人类如何识别高级文本到图像(T2I)模型在涉及复杂构图因素时的缺陷。研究人员手动选择了651张具有复杂构图特征的参考图像,并手动编辑ChatGPT生成的提示来强调构图因素。然后,他们使用三个选定的T2I模型生成AI图像,并对这些图像进行了全面的主体研究以识别其缺陷。实验结果表明,在CO-AID数据集上训练深度模型可以预测AI生成图像中的缺陷并优化AI图像生成,证明了其可用性和有效性。数据集和补充材料可在GitHub上获取。

原文 · arXiv cs.AI

When Composition Doesn't Add Up: Humans Identifying Defects in AI-Generated Images

*Chulin Zhao and Ruoqi Hu contributed equally to this work. State-of-the-art text-to-image (T2I) models exhibit pronounced and systematic defects when prompts involve intricate compositional factors such as multiple entities and multiple attributes. In this paper, we investigate how humans identify such defects. Specifically, we manually select 651 reference images from the four categories of people, hand, object, and scene that exhibit complex compositional characteristics, from which prompts emphasizing compositional factors are derived by manually editing ChatGPT-generated prompts. We then feed the prompts into three selected T2I models to generate AI images and conduct a comprehensive subjective study to identify their defects. For each image, 29 participants provide multi-label assessments specifying defect types and locations. The study yields the compositional AI-generated image defect (CO-AID) dataset, including reference images, prompts, AI-generated images, and information on defect locations and types. Experimental results show that training a deep model on CO-AID can both predict defects in AI-generated images and optimize AI image generation, demonstrating its usability and effectiveness. The database and supplementary materials are available at: https://github.com/Future-IQA/CO-AID .