想看看AI到底能不能做会计?新出的APEX-Accounting基准拿160道真实会计题考了九个模型,最强Claude Fable-5才56%,离实际干活还差得远。
Mercor联手Ramp发布APEX-Accounting,内含160个会计任务,覆盖对账、费用计提等真实工作。九个前沿模型参与测试,Claude-Fable-5 (Max)以56.4% Mean Criteria@3领先Muse-Spark-1.1 (xHigh)的52.6%。最高Pass@8仅21.5%(Muse-Spark-1.1),GPT-5.6-Sol (Max+Pro)的Pass^8为2.6%。实验发现辛普森悖论:提高token预算提升总分,但预算受限时高token消耗任务得分更低。
APEX-Accounting
We introduce APEX-Accounting, a benchmark built by Mercor in partnership with Ramp, to assess whether frontier models can do the real work of accountants. Tasks include reconciling accounts, accruing expenses, posting transactions, and producing reports. The private eval set comprises 160 tasks, split across 10 worlds. Each world contains an accounting system, as well as spreadsheets, PDFs, and other files. Every task was authored and solved by experts in accounting and bookkeeping, who also wrote grading rubrics. Across nine frontier models, Claude-Fable-5 (Max) leads with 56.4% Mean Criteria@3, ahead of Muse-Spark-1.1 (xHigh) at 52.6%. No model scores more than 2.6% Pass^8 (GPT-5.6-Sol (Max+Pro)) and the highest Pass@8 is 21.5% (Muse-Spark-1.1 (xHigh)). We experiment with increasing the token budget from $1 to $50 and observe an instance of Simpson's paradox: scores increase as the token budget increases but within a given budget-constrained harness, scores are lower on tasks where the model spends more tokens. As APEX-Accounting is a closed benchmark, leaderboard evals can be run for any frontier model on request.