研究者提出LLM自我建模评估基准,用合成数据提升模型自我认知能力,但不构成真正的内省。
研究团队评估了LLM回答自身行为问题的能力,引入了可验证行为问题的基准测试。当前模型展现出非平凡但有限的自我建模技能,在关于自身行为的简单反事实问题上存在系统性错误。研究团队开发了可扩展的合成数据管道生成自我建模训练数据,并通过强化学习提升了三个开源模型家族的总体自我建模技能,部分技能可迁移到保留任务上。
Evaluating and Improving LLM Self-Modeling
We study self-modeling: an LLM's ability to answer questions about its own behavior. We focus on verifiable behavioral questions, such as whether a prompt edit would change the model's final answer. To measure this capability, we introduce a benchmark that tests diverse types of self-modeling questions. Current models show non-trivial but limited self-modeling skill, and make systematic mistakes on simple counterfactual questions about their own behavior. To improve self-modeling skill, we develop a scalable synthetic-data pipeline that produces self-modeling training data, and show that reinforcement-learning can improve aggregate self-modeling skill across three open-source model families with some transfer to held-out tasks. These gains, however, do not seem to constitute introspection consistently: improved self-modeling may not arise from privileged access to the model's internal decision process.