BazaarBench:评估 LLM 代理在 C2C 市场中的交易安全
BazaarBench: Delegation Safety in Decentralized C2C Marketplaces Run by LLM Agents
五家模型都能把一件货同时卖给多个买家,对抗指令下更是变本加厉,GPT-5.4 违约率到 55.5%。想做交易代理的先看看这套基准。
研究团队发布 BazaarBench,一个模拟 C2C 二手交易市场的基准,用于评估 LLM 代理替用户交易时的安全风险。基准运行 3 个市场各 30 个模拟日,每个市场 100 个代理,库存来自公开的 eBay 样本,并通过记录核查与 LLM 评分识别 6 类失败。在 45 次延续测试中评估 5 个模型,结果所有模型在普通指令下都会尝试把同一件商品承诺给多个买家。在对抗性指令下,测试卖家完成不可用或夸大商品交易的比例从 15.4% 升至 33.4%,其中 GPT-5.4 达到 55.5%;单个代理模拟周收入也从 20 美元升至 33 美元。团队开源了模拟器、市场状态、评估代码及 357,608 次模型调用记录。
BazaarBench: Delegation Safety in Decentralized C2C Marketplaces Run by LLM Agents
In decentralized consumer-to-consumer (C2C) marketplaces, people list goods, negotiate with strangers, and rate one another, so trust rests on reputation. Large language model (LLM) agents now act for users, raising risks to their money, privacy, and reputation. We introduce BazaarBench, a simulated C2C marketplace and benchmark for evaluating the safety of these agents. It tracks ownership, item condition, and commitments across transactions, combining record checks with rubric-based LLM judgments to identify six failure types across five stages. We run three base markets for 30 simulated days, each with 100 agents using one model and inventories drawn from a public eBay sample. Across 45 continuations, we evaluate five models under ordinary instructions, deadline pressure, or adversarial instructions to exploit other traders. Each continuation runs for seven simulated days from a copy of a market's day-30 state. The tested model controls the same 20 selected agents, retaining their personas, inventories, and histories, while the other 80 keep the base model. All five models attempt to promise the same item to multiple buyers under ordinary instructions. Adding targets and deadlines increases these attempts for every model. Under adversarial instructions, the share of tested sellers' committed transactions completed despite unavailable items or overstated conditions rises from 15.4% to 33.4%, reaching 55.5% for GPT-5.4. Averaged across models and markets, simulated weekly earnings per tested agent rise from USD 20 under ordinary instructions to USD 33 under adversarial instructions. Most of the increase comes from items the sellers never held. We release the simulator, saved market states, evaluation code, and records covering 357,608 agent model calls for evaluating new models and developing safer marketplace agents.