论文精选

大模型代理商业评估中的有效性验证失败

When Guardrails Look Effective: Construct Validity Failures in LLM Agent Commerce Evaluation

精选理由

这篇论文揭示了AI代理评估中常见的验证漏洞,教你如何识别虚假有效的市场护栏。

AI 摘要

研究人员在酒店交易测试中发现,LLM代理评估中的市场护栏效果存在严重验证问题。初始报告显示Qwen2.5 1.5B-14B模型带来+87.4、+35.0和+28.8的福利增益,但固定架构后变为+7.2、-13.9和+23.8。14B模型单次生成平均效应为+229,三次生成后降至+37.6,生成残差解释了49.9%的变异。

原文 · arXiv cs.AI

When Guardrails Look Effective: Construct Validity Failures in LLM Agent Commerce Evaluation

Interactive simulations increasingly evaluate policies in markets populated by language-model agents. Their outputs can look economic---prices, profits, consumer surplus, and welfare---without instantiating the behavior named in the claim. We audit this risk in a multi-turn buyer--seller testbed for configurable hotel transactions. An initial implementation reported welfare gains from two marketplace guardrails of +87.4, +35.0, and +28.8 across a Qwen2.5 1.5B--14B ladder. It also gave guarded and unguarded agents different offer schemas and choice procedures. Holding the schema and buyer chooser fixed changes the paired contrasts to +7.2, -13.9, and +23.8. The four largest 14B single-generation effects averaged +229; after three generations per profile-condition, they averaged +37.6 (95% bootstrap interval [-34.2, 109.3]), while generation residuals account for 49.9% of variation in this post-hoc probe. A seller-incentive check is non-monotone: increasing profit pressure produces less profit than the default seller prompt. Scripted positive controls show why this matters. A profit-maximizing seller already attains first-best welfare, so guardrails mostly redistribute and reduce welfare; they create welfare only when the seller is explicitly programmed to force inefficient bundles. We contribute a construct-validity contract separating incentive validity, protocol isolation, stochastic stability, and welfare accounting, and returning INVALID or INCONCLUSIVE before substantive policy claims. In our case, the original estimate is INVALID under protocol isolation, while the controlled study remains INCONCLUSIVE under incentive validity and stochastic stability. The case does not show that guardrails are ineffective; it shows their apparent value is unidentified until the simulated agents and protocol pass these checks.

大模型代理商业评估中的有效性验证失败 · AI 热点