想知道AI能不能当靠谱的社交裁判?这篇用34个模型和198个人类实测对比,结论挺有意思。
该论文检验了LLM作为社交评判者的有效性,基于10个心理和关系构念构建了3层人格画像(社交吸引、混合、不吸引)。研究1中,34个LLM对12个画像进行3次重复评分,表现出高稳定性和一致的3层排序。研究2通过6对匹配姓名和代词及仅代词测试,未发现性别呈现的显著影响。研究3中198名人类参与者的评分复现了3层结构,但LLM对吸引画像评分更积极、对不吸引画像更消极。
How Do Large Language Models Judge Social Attraction? Evidence from Theory-Grounded Persona Ratings Across Multiple LLMs and Humans
Large language models (LLMs) are increasingly used to perform subjective evaluations traditionally made by humans, yet their validity as social judges remains unclear. This paper examines whether LLMs can assess social attraction from theory-grounded persona profiles constructed from ten psychological and relational constructs and organized into three tiers: socially attractive, socially mixed, and socially unattractive. We examine LLM ratings in two studies and compare them with human judgments in a third study. In Study 1, 34 LLMs rated 12 profiles across three repeated runs. Although some models tended to give higher or lower ratings overall, they showed strong stability across runs, consistent three-tier ordering, and high agreement in relative profile ordering. Study 2 examined sensitivity to gender presentation using six matched name-and-pronoun profile pairs and a separate pronoun-only test with a gender-neutral name, finding no significant effects in either analysis. In Study 3, 198 human participants evaluated the six matched profiles from Study 2. Their ratings reproduced the three-tier structure and followed a profile ordering consistent with that of the LLMs. However, LLMs rated attractive profiles more positively and unattractive profiles more negatively than humans, while neither group showed a significant overall effect of gender presentation.