技巧

模型与工具适配性研究

Finding the Right Fit: Model-Harness Interactions across Agent Tasks

精选理由

研究分析了66种模型与工具组合,发现高成本不一定带来高分,不同模型在不同工具中表现差异显著。

研究评估了66种配置,包括四种工具适配器(OpenHands、DeepSeek Harness、PI和openJiuwen)与五种模型在TUA-Bench、ALE-CLI和Terminal-Bench 4基准测试中的表现。模型在不同工具适配器中的排名会发生变化,例如在Terminal-Bench 4上,Claude在OpenHands中领先GPT 7.94分,但在PI中落后30.16分。openJiuwen为Kimi在所有三个基准测试中提供了最高分,领先5.61至11.11分。

原文 · arXiv: DeepSeek

Finding the Right Fit: Model-Harness Interactions across Agent Tasks

Choosing an agent system means choosing both a language model and the harness through which it acts. We ask whether a strong model, harness, or pairing stays strong when the setting changes. We evaluate 66 configurations: four configurable harnesses (OpenHands, DeepSeek Harness, PI, and openJiuwen) paired with five models on TUA-Bench, ALE-CLI, and Terminal-Bench 4, plus the native Codex-GPT and Claude Code-Claude pairings. Model rankings reverse across harnesses. On Terminal-Bench 4, Claude leads GPT by 7.94 points in OpenHands but trails it by 30.16 points in PI. For four of the five models, the best harness changes from one benchmark to another, yet some pairings hold: openJiuwen gives Kimi its highest score on all three benchmarks, by 5.61 to 11.11 points. A model's own vendor harness is not reliably its best, and higher cost does not reliably buy a higher score. On Terminal-Bench 4, GPT scores higher under PI than under DSH at less than a quarter of the cost per task. Matched trajectories suggest why fit varies. Models start almost all repairs themselves, so much depends on whether the harness hands failures back in a form the model can use. GPT does best with PI's lean scaffold, while Kimi, which often issues malformed tool calls, does best in openJiuwen. We argue that the model, the harness, and the task should be evaluated together, and we release the harness adapters, evaluation code, and all 6,204 scored trajectories at https://github.com/liyix/finding-the-right-fit and https://huggingface.co/datasets/yixuanli97/finding-the-right-fit.