研究评测 DeepSeek、Qwen、Doubao 的谄媚行为:36 万条回答分析
Evaluating Sycophancy in Chinese Large Language Models on Factual Questions Derived from Online Search Queries
研究者拿 12 万多个中文事实问题测了 DeepSeek、Qwen 和 Doubao,发现让模型别顺着你说话,答案反而会变得含糊。做中文产品的话值得看看。
一项研究用 12,165 个来自真实中文搜索查询的事实判断题,分析三个中国前沿模型 DeepSeek、Qwen 和 Doubao 的谄媚行为,共收集 364,941 条回答。实验对比了基线、用户信念条件与反谄媚提示三种设置,并区分两种失败模式:错误信念引发的错误回答,以及原本正确的回答变得不确定。结果显示推理模式并非稳定保障,反谄媚指令能减少对错误信念的附和,但会增加回答的不确定性。研究表明在中文事实问答中,不附和错误信念不等于保持事实准确性。
Evaluating Sycophancy in Chinese Large Language Models on Factual Questions Derived from Online Search Queries
As large language models increasingly mediate information access, factually accurate and independent answers are critical. However, these models can exhibit sycophancy by aligning their responses with users' stated beliefs even when those beliefs are incorrect, potentially presenting misinformation as independently verified and reinforcing users' confidence in false claims. Prior work leaves unresolved whether introducing user beliefs causes correct responses to become incorrect or uncertain, or causes uncertain responses to become belief-aligned incorrect answers. It also remains unclear whether anti-sycophancy interventions preserve or restore factual accuracy or merely shift responses toward uncertainty. We analyze factual sycophancy in Chinese-language information seeking using yes/no fact-checking questions. Our analysis covers 364,941 responses from three frontier Chinese-based LLMs (DeepSeek, Qwen, and Doubao) to 12,165 factual questions derived from real-world Chinese search queries. We evaluate the models with and without reasoning across baseline, belief-conditioned, and anti-sycophancy prompting, tracing matched shifts among correct, incorrect, and uncertain responses. Under incorrect user beliefs, we distinguish belief-aligned errors from losses of factual confidence, in which initially correct answers become uncertain. Patterns vary across models and reasoning settings: reasoning is not a consistent safeguard, and anti-sycophancy instructions can reduce incorrect agreement while increasing uncertainty. In Chinese-language factual question answering, avoiding agreement with false beliefs is therefore not equivalent to preserving factual accuracy, highlighting the value of transition-level evaluation. Such behavior may undermine the reliability of LLM-mediated information access by reinforcing misinformation or weakening users' confidence in factually correct answers.