论文

研究:harness自进化在DeepPlanning任务上收益有限

精选理由

有人拿DeepPlanning这个多日旅行规划基准来讨论:harness自己改自己,到哪一步就不管用了。做agent评测的可以看看这个思路。

一条推文引用Zhang等人2026b的DeepPlanning基准,该基准要求多日旅行规划任务调用航班、酒店、景点、距离等工具。任务需要跨多步追踪状态,并输出结构化计划,由程序按硬约束打分。推文据此讨论harness自进化(如DSH creator mode或prime-agent)在何种任务上不再带来性能提升。

原文 · Teortaxes

Related: we can reasonably determine where harness self-evolution (a la DSH creator mode or prime-agent, perhaps) becomes insufficient for performance gain on a task. Hill to climb: «DeepPlanning (Zhang et al., 2026b) poses multi-day travel-planning tasks that require tool calls (flights, hotels, attractions, distances), state tracking across many steps, and a final structured plan that is scored programmatically against hard constraints.»