论文精选73°

伯克利论文:LLM知道偏好变化仍使用旧值

精选理由

伯克利发现LLM知道用户偏好变化却仍用旧值,解决方法很简单:在提示中明确当前状态。

伯克利研究显示,当用户偏好或截止日期变化时,旧版本信息仍留在上下文中。5个开源模型测试表明,将注意力引导到最新值可修复大部分错误。GPT-5.6 Sol在长代理日志测试中仅答对9/40题,但提供当前状态后全部答对。

原文 · rohanpaul_ai

New Berkley paper: LLMs often know you changed your mind but still use your old choice, so agents need the current state spelled out.

When a preference or deadline changes, the old version stays in context. The model still holds the new one, but its attention keeps drifting back to older mentions.

In 5 open models, nudging attention toward the newest value fixed most of these mistakes without retraining. Even top-tier GPT-5.6 Sol got only 9 of 40 questions right on long agent logs, but 40 of 40 when given the current state.

If your agent tracks anything that changes, keep the current state in the prompt instead of making the model dig through history.

– arxiv. org/abs/2609.38866

Title: "When Context Changes: Understanding Update Failures in LLMs"