论文

Argo-Bench评估数据代理企业级工作流

Argo-Bench: Evaluating Data Agents on Enterprise-Scale Workflows

精选理由

Argo-Bench用真实规模数据测试AI代理,发现顶级模型完成率不足35%,暴露了当前AI在企业数据环境中的局限性。

Argo-Bench是一个包含210个数据科学和分析任务的评估框架。它模拟了纽约市一个拥有8100万份订单的食品配送平台,构建了包含235个表和75亿行的ERP仓库。14个前沿和开源模型中表现最好的模型仅在34.8%的任务上得分95或以上,平均得分为59.5分。

原文 · arXiv cs.AI

Argo-Bench: Evaluating Data Agents on Enterprise-Scale Workflows

Real-world enterprise data science and analytics workflows require reasoning across dozens of tables, performing statistical analyses, and acting on the results. Established text-to-SQL benchmarks evaluate query generation alone, and audits have found their answer keys frequently wrong. Because real enterprise warehouses are too sensitive to release, these benchmarks are built on public datasets where a business event fits in a single table. We introduce Argo-Bench, an evaluation framework comprising 210 data science and analytics tasks. Drawing on public data, peer-reviewed industry literature, and regulatory filings, we simulate a food delivery platform in New York City at true scale, with 81 million orders in 2024, grounded economics, fraud patterns, and marketplace incentives. We export this world to an ERP warehouse of 235 tables and 7.5 billion rows, modeled on the Oracle E-Business Suite schema. The simulator's ground-truth state is withheld from the warehouse the agent sees, so tasks require reconstructing facts by navigating the warehouse before acting on them. Argo-Bench goes beyond text-to-SQL: the agent files actions such as banning fraudulent accounts, allocating courier incentive budgets, or issuing back pay, and the grader scores each by its consequences in the simulator. Every task has an executable reference solution that demonstrates solvability using only the warehouse. The strongest of 14 frontier and open-weight models scores 95 or higher on only 34.8% of tasks and averages 59.5 points. We hope Argo-Bench drives progress toward agents that understand, navigate, and act within real data environments.