论文多源确认精选

Stanford 论文:智能体失败先分清流程与内容问题,再决定改 harness 还是微调

精选理由

Stanford 这篇研究很实用:智能体死循环就改 prompts 和工具,计划烂就上 LoRA 微调,还给了一组具体数字对比。

Stanford 的一项研究把智能体失败分为流程失败(死循环、步数耗尽)和内容失败(给出的计划质量差)两类,两者对应不同的修复手段。在旅行规划基准上,LLM 驱动的 harness 进化把 Qwen3.5-4B 的 held-out 成绩从 0.16 提到 0.30,计划交付率从 55% 升到 90%。但 harness 编辑从未降低差计划占比,而 LoRA 适配器把 Qwen3.5-9B 差计划占比从 28% 降到 5%。结论是:先标注运行失败的原因,流程问题改 harness,内容问题动权重。

原文 · rohanpaul_ai

New Stanford paper finds that harness changes fix agents that loop or stall, while agents that deliver bad plans need weight training instead.

An agent can be improved by editing its harness, the prompts, tools, and checks around the model, or by fine-tuning its weights.

They sorted failed runs into process failures, such as loops and used-up step budgets, and content failures, where a poor plan was delivered. On a travel-planning benchmark, an LLM-driven loop rewrote the harness, and its best runs then fine-tuned the model.

Harness evolution lifted Qwen3.5-4B from 0.16 to 0.30 on held-out tasks, as plan delivery rose from 55% to 90%. Harness edits never shrank the share of poor plans, but a LoRA adapter cut them from 28% to 5% of Qwen3.5-9B's held-out runs.

Before improving an agent, label why its runs fail, then fix process failures in the harness and content failures in the weights.