论文精选

研究显示大语言模型推理能力提升心理深度需看评判者是谁

Does Reasoning Improve Psychological Depth in Large Language Models? It Depends on Who's Judging

精选理由

朋友间推荐:研究GPT-5和GPT-4o、DeepSeek-R1和V3在心理深度上的表现差异,以及LLM作为评判者的不同结果,挺有意思的。

研究通过心理深度短篇故事测试,发现人类读者对GPT-5和GPT-4o的偏好无统一结论(GPT-5略占优),但对DeepSeek-R1和V3则相反(V3更优)。而LLM作为评判者时,在89%的维度比较中偏好推理输出。研究指出开发集表现不足以证明部署有效性。

原文 · arXiv: DeepSeek

Does Reasoning Improve Psychological Depth in Large Language Models? It Depends on Who's Judging

LLM-as-a-Judge evaluators are increasingly used to score open-ended generation, yet a judge's correlation with human ratings on its development set may not guarantee valid measurement when outputs are closely matched and human preferences are subjective. We study this failure mode through psychological depth in short stories. Seven human readers and an LLM-judge ensemble selected on the original scalar Psychological Depth Scale dataset ($ρ= 0.646$) evaluated 60 blinded, prompt-matched story pairs from GPT-5 vs.\ GPT-4o and DeepSeek-R1 vs.\ DeepSeek-V3. Human preferences showed no universal reasoning advantage: GPT-5 was modestly preferred over GPT-4o (60.0--62.9\%), whereas DeepSeek-R1 trailed V3 (42.9\%), and inter-reader agreement was near chance (Krippendorff's $α= 0.070$), with within-reader consistency and recurring weighting patterns suggesting structured heterogeneity rather than random responding. The judge, by contrast, favored reasoning outputs in 89.0\% of dimension-level comparisons and 59 of 60 pairs on aggregate PDS, uniformly across all five evaluator configurations, and its scores were associated with surface features such as sentence length and lexical diversity. These results suggest that development-set performance is insufficient evidence for deployment validity on a shifted distribution, and that point-estimate judges can obscure the heterogeneity in subjective human evaluation.