技巧

软件工程代理中的工具链效应研究

Beyond the Model: Demystifying Harness Effects in Software Engineering Agents

精选理由

论文分析了不同工具链组件对编程助手性能的影响,帮你理解如何构建更高效的AI编程助手。

该研究评估了两个代表性工具链mini-SWE-agent和OpenCode在Qwen和DeepSeek模型上的表现,使用SWE-bench Pro、ProgramBench和GitTaskBench三个基准测试。研究构建了NanoHarness轻量级模块化工具链,分析了五个关键组件的影响。实验表明,工具链效果取决于模型能力和任务类型,复杂工具链在模型能力提升时对SWE风格问题修复的边际收益递减,但在更复杂的仓库级任务中能增强强模型性能。

原文 · arXiv: DeepSeek

Beyond the Model: Demystifying Harness Effects in Software Engineering Agents

Large Language Model (LLM)-based agents are increasingly used for software engineering tasks, yet their performance is not determined by the base model alone. The agent harness substantially shapes how SE agents interact with repositories, execute actions, and validate solutions. However, the role of harness design remains insufficiently understood, especially across different models, tasks, and harness components. In this paper, we present a systematic empirical study of harness effects in SE agents. We first evaluate two representative harnesses, mini-SWE-agent and OpenCode, with ten models from two prominent open-weight model families, Qwen and DeepSeek, on three benchmarks: SWE-bench Pro, ProgramBench, and GitTaskBench. We then construct NanoHarness, a lightweight modular harness built on top of mini-SWE-agent, and use it to analyze five representative harness components: tool registry, context compression, explicit planning, subagents, and lazy skills. Experimental results show that harness effectiveness depends jointly on model capability and task type. Complex harnesses provide diminishing marginal gains on SWE-style issue repair as model capability improves, but can benefit stronger models on more complex and open-ended repository-level tasks. Component-level analysis on ProgramBench further shows that structured tool use and task-specific subagents provide the most stable improvements, while context compression and general subagents can hurt repository-generation performance. When combined, NanoHarness improves over mini-SWE-agent by 7.37 and 6.21 percentage points on Qwen3.7-Max and DeepSeek-V4-Pro, respectively, recovering most of the gains of product-level harnesses. These findings highlight harness design as a first-class factor in SE-agent performance and provide insights for building more effective and efficient coding agents.