论文精选

研究表明:换 Agent 框架对结果的影响堪比重跑一次

精选理由

同一模型跑 Claude Code、mini-SWE-agent、OpenCode 对比,结论很实用:精简提示词和工具能省最多 3 倍成本,写 Agent 的都该看看。

该研究用同一个模型分别跑了 Claude Code、mini-SWE-agent 和 OpenCode 三种 Agent 框架,在 SWE-bench Verified 的 447 个任务上对比。Claude Code 与 mini-SWE-agent 的成绩差距在 5 分以内。在 45 个高难度任务上,更换框架与重复运行同一框架导致的任务翻转比例均为 13%。成本差异主要来自系统提示词和工具 schema——每个框架每一步都重复发送这些内容,步数越多花费越高。研究给出的建议是把框架提示词和工具精简到最小,可最多降低 3 倍成本且不损失准确率。

原文 · DAIR.AI

Useful paper on what an agent harness changes when the model stays the same.

One takeaway: keep your harness prompt and tools small.

In this study, that can cut costs by up to 3x without lowering accuracy.

The authors ran Claude Code, mini-SWE-agent, and OpenCode with the same model on SWE-bench Verified. Claude Code and mini-SWE-agent scored within 5 points of each other on 447 tasks.

Swapping the harness changed results about as much as rerunning it. On a 45-task hard set, both flipped 13% of tasks.

The cost difference comes from the system prompt and tool schemas, which each harness sends again at every step. The more steps the agent takes, the more times you pay for them.

Paper: https://t.co/pwsrCws67u