新基准用《秘密希特勒》测AI说谎,16个模型里GPT-5.4、Kimi K2.5领先。
ParliamentBench 是基于桌游《秘密希特勒》的开源评估框架,用于测试大模型在信息不对称下的欺骗、说服与推理能力。研究对16个LLM进行了1600场模拟对局,并与人类玩家及在线游戏数据对比。排名前四的模型为 GPT-5.4、Kimi K2.5、Grok 4.1 Fast 和 DeepSeek 3.1 Terminus,而最弱模型的表现低于随机基线33%和算法基线45%。研究还引入了三项新指标,分别衡量社交推理、推理能力和欺骗一致性。多数模型难以维持一致的欺骗人格,欺骗保持率低于50%。
Can Agents Deceive? Evaluating Reasoning and Deception in ParliamentBench using a Social Deduction Game
As large language models (LLMs) are deployed as agents in high-stakes settings, such as medical and legal systems, understanding their deceptive capabilities is fundamental to safety. Controlled social deduction games provide a reproducible proxy for isolating and evaluating these complex adversarial behaviors. We present the open-source benchmark framework ParliamentBench based on the game Secret Hitler to evaluate LLMs in scenarios that require deception, persuasion, and reasoning under information asymmetry. We evaluate 16 LLMs across 1,600 simulated matches playing each other, playing against humans, and compare them against a large set of online games. We introduce three novel metrics that isolate social deduction, reasoning, and deceptive consistency. Our experiments reveal that frontier models achieve strong performance across cooperative and deceptive roles, with a strong top-four cluster (GPT-5.4, Kimi K2.5, Grok 4.1 Fast, and DeepSeek 3.1 Terminus), whereas the weakest models fall short of random (33%) and simple algorithmic (45%) baselines. Most LLMs struggle to maintain a consistent deceptive persona throughout an entire game, with deception retention dropping below 50%.