LMArena 分析智能体错误归因:GPT-6 Luna 误归因率达 53.1%
LMArena 拆了个挺冷门的问题:智能体会把话安到你头上。GPT-6 Luna 误归因率 53.1%,看看各家模型差在哪。
LMArena 发布了针对智能体 False Attribution(错误归因)行为的分析,指智能体将陈述、请求或事实错误地归于用户,与用户提供的证据相矛盾。GPT-6 Luna 和 Astra 很少误引用户原话(分别 15.6% 和 28.6%),但误归因率高(53.1% 和 48.2%)。它们的同系列模型 GPT-6 Sol 误述用户历史 rate 最高,达 23.5%。不同模型在误引与误归因两类错误上呈现不同模式。
We also looked into False Attribution, where the agent attributes a statement, request, choice, approval, or fact to the user, but user-provided evidence contradicts that attribution. Interestingly we saw some models misquoting the user (misstating what a user asked for), while others credit the user with someone else's work. There were varied patterns across models, with GPT-6 Luna and Astra rarely misquoting (15.6% and 28.6% respectively) but often misattributing (53.1% and 48.2%). Interestingly their sibling model GPT-6 Sol has the highest rate of misstaging the user’s history (23.5%). 💬 1 🔄 0 ❤️ 3 👀 668 📊 1 ⚡