研究:强智能体在ML工程中需要多少约束
How Much of a Harness Does a Strong Agent Need for Autonomous ML Engineering?
论文发现复杂约束系统在MLE任务中不如简单编码基线,LLM backbone才是性能主要驱动因素。
该论文研究了自主机器学习工程(MLE)智能体在相同时间预算和前沿LLM backbone下,复杂约束系统与最小约束编码基线的性能对比。研究通过大规模系统性消融实验发现,开源的最先进约束系统在MLE基准测试中并未展现出优势。结果表明,当前MLE基准测试中,围绕强模型精心设计的约束层收益甚微。
How Much of a Harness Does a Strong Agent Need for Autonomous ML Engineering?
Recent autonomous machine learning engineering (MLE) agents have made significant progress on public leaderboards. Often motivated by progress stagnation over long-horizon cycles and limited Large Language Model (LLM) primitives, modern MLE agents are deployed on top of increasingly elaborate machinery: multi-agent orchestrators, dedicated retrieval subagents, and more. While such harnesses expand, the use of more primitive but improved coding agents - where LLMs have direct access to the execution environment through read, write, and bash primitives - has received little attention in the field. In this paper we find that, under an equal time budget and the same frontier LLM backbone, open-source state-of-the-art harnesses provide no advantages over a single session of a minimal-harness coding agent baseline, pointing to the backbone as the primary driver for performance. Via a series of large-scale systematic ablation studies, we argue that the machinery layers become redundant in the coding agent setting. We conclude that the effort spent elaborating hand-crafted harnesses around strong models yields poor returns for current MLE benchmarks.
- Apple ML Research10-01 00:00原文