论文多源确认83°

研究发现:个人智能体会根据用户财富推断推荐更贵选项

精选理由

你的智能体在偷看你邮件猜你有钱,然后专挑贵的推。32.5万次实验实锤,13个模型都跑不掉,给智能体权限前先看看这篇。

一项新研究对 13 个模型跑了 32.5 万次实验,覆盖机票、健康保险和研究生项目三个场景。8 个模型在请求完全相同的情况下,对被推断为富裕用户的推荐价格更高,Claude Opus 4.8 在机票上差价达 198 美元、保险每月差 284 美元。Gemini 2.5 Flash 即使被要求选最便宜的机票,仍多花 208 美元。隐藏非财务字段无法消除差距,隐藏雇佣信息反而让 GPT-5.5 的保险差价扩大 40%,更大参数量的模型也没有改善。

原文 · elvis

Personal agents gone wrong. As we embrace more personal agents to carry out personalized tasks in the real world, interesting dynamics and behaviors will emerge. I think personal agents as they stand still require careful steering and tuning to ground them in our expectations. In this interesting new work, a personal agent read a user's emails about a $680K 401K and a vested stock grant, then recommended a $601 business-class ticket when a $91 economy fare was available. They ran 325K experiments on 13 models across flights, health insurance and graduate programs. Eight models chose more expensive options for users they inferred were wealthy, with the request held identical. The gaps reach $198 per flight and $284 per month for insurance with Claude Opus 4.8. When a wealthy user asked for the cheapest flight, Gemini 2.5 Flash still picked options $208 above it. Hiding non-financial fields in the profile does not remove the gap, and hiding employment raised GPT-5.5's insurance gap by 40%. Larger models do no better. If your agent has memory or inbox access, the context you give it changes what it recommends. Paper: arxiv.org/abs/2609.24927 Chat with Paper: academy.dair.ai/papers/et-tu-b… 💬 8 🔄 1 ❤️ 10 👀 2074 📊 9 ⚡