LLM作为调查替代品的品味偏差研究:系统正偏与结构缺失

Not-quite-human tastes: the stylized omnivorousness of LLM survey surrogates

精选理由

这篇论文揭示了用AI模拟人类文化品味时的三个致命缺陷:过度喜欢、关系缺失和社会偏差。做市场调研的人最好先读一读。

AI 摘要

该研究使用OpenAI、Anthropic和DeepSeek的LLM为每个模型生成277,470个(30×9249)硅样本,基于美国艺术参与调查(SPPA)数据。研究发现硅样本对喜好存在系统性正偏差,使生态估计值膨胀;样本间的关系结构完全丢失;年龄-品味关联被削弱,阶级-品味关联被复活,性别和种族-品味关联被夸大。

原文 · arXiv: OpenAI

Not-quite-human tastes: the stylized omnivorousness of LLM survey surrogates

Large-language models have proven to be remarkable if inconsistent parrots of public attitudes and opinions. The extent to which LLMs are able to produce reasonable approximations of cultural taste remains an open empirical question that becomes more urgent by the day, with market research companies already offering provisional `synthetic' survey panels and the contamination of standard survey data from LLM-generated responses. In this study, we build on past work on silicon sampling by extending considerations of its algorithmic fidelity and alignment to the domain of cultural consumption. We use large-language models from OpenAI, Anthropic, and DeepSeek to each produce 277,470 (30x9249) silicon surrogates of survey respondents from the Survey of Public Participation in the Arts (SPPA). We find these silicon surrogates' tastes to be highly stylized facsimiles of human tastes. (1) Silicon samples have a systematic postive-bias for liking, resulting in inflated ecological estimates of tastes. The individual-level bias of silicon samples are not well-explained by the WEIRD-bias often discussed in the literature. (2) The complex relationality in real taste structures is completely lost among silicon samples. (3) Finally, very little of the known cultural alignment between tastes and social space are preserved. Silicon samples attenuate age-taste associations, resurrect anachronistic class-taste associations, caricaturize gender- and race-taste associations.