PriceBench:用酒店预订任务测量 28 个 LLM 的价格与品牌偏好
PriceBench: A Diagnostic Benchmark for Price, Quality, and Brand Preferences in LLM Booking Agents
同一批酒店任务,28 个模型选出的房价能差 150 美元,这篇论文把这个差异量化出来了,做购物或预订代理的人应该看看。
研究团队发布 PriceBench 基准,用 logit 选择模型从预订选择中还原 LLM 的价格、质量和品牌偏好,覆盖 8 家提供商的 28 个 LLM,任务来自 179 家纽约真实酒店的 3,600 个酒店预订场景。结果显示能力与偏好强度相关而与偏好内容无关:更强的模型偏好更一致,弱模型要么锁定单一选项、易被列表顺序操纵,要么近乎随机。价格敏感度跨提供商相差超过一个数量级,同等任务下平均每晚成交价从 247 美元到 393 美元不等。论文指出代理买什么必须逐模型实测而非推断,任务、代码和 28 份响应数据均已开源。
PriceBench: A Diagnostic Benchmark for Price, Quality, and Brand Preferences in LLM Booking Agents
LLMs increasingly act as purchasing agents, which makes the LLM, not the user, the one choosing among the options that satisfy a request; its preferences quietly fix what gets bought and what it costs. Hotel booking is a clean instance: a high-volume choice settled on a few comparable attributes, where the pick reveals those preferences. We introduce PriceBench, a diagnostic benchmark that recovers an LLM's price, quality, and brand preferences from its booking choices with a logit choice model, applied to 28 LLMs from 8 providers on 3,600 hotel tasks from 179 real New York City properties. We find that capability is associated with how consistently an LLM chooses, not with what it chooses: more capable LLMs hold stronger, more consistent preferences, while weaker ones either lock onto one position, exploitable by whoever controls listing order, or choose almost indifferently. What those preferences favor varies sharply across providers and even within one family: price sensitivity spans more than an order of magnitude, and the price/quality trade-off moves mean booked nightly price from \$247 to \$393 on identical tasks. What an agent buys must therefore be measured per LLM, not inferred, and we release the tasks, code, and all 28 response sets.