论文精选73°

研究评估人类与大语言模型的心智化能力

Assessing mentalization in humans and large language models

精选理由

这篇论文用经济游戏测试了四大模型家族的心智化能力,发现GPT-5能灵活调整推理深度,表现超过人类。

AI 摘要

研究人员通过两个经济游戏和认知计算模型,测试了DeepSeek、GPT-4.1、GPT-5和Gemini 2.0 Flash共2099个LLM代理的心智化能力。结果显示,不同模型提供商和大小的LLM表现出明显不同的心智化行为和计算特征。GPT-5代理能够灵活调整其递归推理深度,表现优于251名人类参与者。

原文 · arXiv: DeepSeek

Assessing mentalization in humans and large language models

Mentalization - the ability to infer others' beliefs and intentions to guide one's own choices - is a key cognitive function underlying human social interactions. Large language models (LLMs) demonstrate behaviour consistent with humans on theory-of-mind tasks, yet whether these models can guide adaptive behaviour through mentalization is unknown. Here we use two economic games with cognitive computational modeling to uncover the latent strategies underlying mentalization in LLMs. We tested individual LLM agents across four model families, DeepSeek, GPT-4.1, GPT-5 and Gemini 2.0 Flash (N = 2,099), against opponents of varying sophistication and examined whether a prompting strategy designed to elicit strategic reasoning improved performance. We benchmarked results against human participants (N = 251) as a comparative measure. Across both games, LLMs showed clear behavioural and computational signatures of mentalizing that differed markedly by model provider and size. Strategic prompting generally improved performance by inducing more sophisticated reasoning, yet the extent of the benefit differed across the two tasks. Last, GPT-5 agents flexibly adapted their recursive depth of reasoning to increasingly sophisticated opponents, demonstrating superior performance to human participants. Collectively, we demonstrate different capacities for mentalization across LLMs, and highlight cognitive computational modeling as a formal method for assessing comparative intelligence across humans and machines.