研究模拟个人 AI 助手购物市场:采纳者省 8.7 美元,无助手者受影响有限
Who Keeps the Gains from Personal AI Assistants? Seller Adaptation and the Unassisted in a Language-Model Market Simulation
用语言模型搭了个模拟市场,实测 AI 助手砍价到底省多少、卖家怎么应对,连没装助手的人会怎样都算了,数字给得很具体。
一篇 arXiv 论文构建了由语言模型扮演消费者、助手和六家自适应卖家的租房市场模拟,共运行 30 个市场。结果显示,执行型助手在卖家策略冻结时为采纳者每天省 13.7 美元,卖家适应后收回约三分之一,剩余 8.7 美元。此前校准预测无助手消费者会多花 3.6 美元,但实测均值仅 +0.4 美元,置信区间为 -0.6 到 +1.3。论文结论是应从市场层面评估助手,把完成率、总支出和无用户影响都纳入考量。
Who Keeps the Gains from Personal AI Assistants? Seller Adaptation and the Unassisted in a Language-Model Market Simulation
Personal AI assistants are beginning to transact for consumers, and early adopters capture real savings. Whether those savings survive, and what happens to consumers who have no assistant, depends on how sellers respond -- a question single-user evidence cannot answer. We build an agent-based rental market in which language models play consumers, assistants, and six adaptive sellers guided by an algorithmic pricing tool. Half the population receives an assistant under an advisory or an executing mandate; the contract pairs a fee only the renter's physical action avoids with a pre-selected add-on an authorised assistant can cancel online. An analytical benchmark and a behaviourally calibrated rule market supply ex-ante predictions, and paired branches with frozen versus adaptive sellers separate adoption effects from market feedback. Across thirty simulated markets, executing assistants cut adopters' spending by 13.7 USD per renter-day when sellers are frozen; adaptation claws back about a third, leaving 8.7, with the gains arriving both as lower bills and as rentals completed at all. Sellers raise headline rates while cutting fees, and the calibrated forecast of the burden on unassisted consumers (+3.6) does not transfer: their mean spending change is +0.4, confidence interval -0.6 to +1.3. Seller-model swaps and a within-market transfer of fee-setting to the pricing tool show that fee conduct, and with it the division of the gains, is decided on the seller side. Assistants, we conclude, should be evaluated at market level -- completion, total spending, and non-users included -- and the comparison layers locate exactly where a calibrated behavioural forecast fails in a language-model market.