这篇论文告诉你,LLM 当模拟器时可能答对结果但用错理由。作者用94人的防晒霜测试证明了这一点,并提供审计方法。
论文提出一个审计框架,通过94名受访者的防晒霜概念测试,将人类开放理由映射为符号推理状态(正负号表示支持或拒绝)。人类理由能显著提升购买意图的预测准确率。LLM 模拟的理由虽然听起来合理,但多重复概念板内容,而非真实受访者的接受或拒绝路径。该框架提供了一种可解释的测试方法,检查 LLM 模拟器的理由是否与人类证据一致。
Reason-Mediated Behavioral Models for Auditing LLM Social Simulators
Large language models are increasingly used as social simulators, including as synthetic survey respondents. Most evaluations ask whether simulated outcomes resemble human outcomes. We argue that this is necessary but too weak: a simulator can match the final answer while using the wrong rationale-derived reason pattern. We study this problem through a 94-person sunscreen concept test in which each respondent evaluated three product concepts and wrote open-ended rationales. We map those rationales into signed reason states $Z$, where positive signs support adoption and negative signs block it. This gives a practical audit: holding respondent descriptors $D$, category context $K$, and concept treatment $X$ fixed, do human rationale-derived reasons help predict behavior $Y$, and can an LLM simulate the same reason state without seeing the human rationale or outcome? Human rationale-derived reasons substantially improve held-out prediction of purchase intent. LLM-simulated reasons are more brittle: they often sound plausible, but frequently echo the concept board rather than recover the respondent's acceptance or rejection path. The paper contributes an evaluation framework for social simulators. Reason states do not identify natural causal effects by themselves, but they provide an interpretable test of whether a simulator's stated reasons align with human evidence.