新研究:用 Agent 历史预测上下文压缩何时造成损害
Does an Agent's History Tell You When Compaction Will Hurt? A Modest, Bounded Effect on the TRACE Paired-Replay Corpus
做长任务 Agent 的朋友看看这篇:TRACE 用 590 个真实压缩边界回放数据测了压缩时机怎么选,结论是历史行为预测力很弱,数据和方法都开源了。
TRACE 在 AppWorld 上的 590 个压缩边界数据集中,将每次压缩在压缩前上下文和摘要两种条件下回放对比。结果显示压缩前的历史行为对压缩后损害的预测力很弱:预注册的对比检验得到宽零结果,扩展协议的最佳触发器在留存集上 AUROC 仅 0.66,低于同边界回放基准的 0.72。最好的冻结可解释触发器能避开 21% 的有害压缩边界,同时保留 84% 的压缩机会。作者指出该数据集无法回答最佳触发器是否优于固定 token 预算规则,并列出了语料库需要补充的内容。
Does an Agent's History Tell You When Compaction Will Hurt? A Modest, Bounded Effect on the TRACE Paired-Replay Corpus
Many long-horizon agents compact their context on a global rule, usually a token budget, blind to what the agent was doing. We ask whether the agent's recent behaviour predicts when a compaction will hurt. TRACE's public corpus of 590 harness-triggered AppWorld compaction boundaries replays each boundary from a re-executed prefix state under the pre-compaction context and under the summary, and records the burden of the next actions: calls that error or repeat a call already made. We find that pre-boundary history predicts post-compaction harm only weakly. An internally prespecified contrast by prefix placement is a wide null, and the naive "has-written" label behind it turns out to measure trajectory phase. The best extension-protocol trigger reaches held-out AUROC 0.66 (0.64 on the replicate's own label) against a same-boundary replicate of 0.72; the best frozen, interpretable trigger avoids 21% of harmful (positive-burden) boundaries while keeping 84% of compaction opportunities, and exceeds the random-rule expectation on count but not on burden mass (a post hoc comparison). Whether the best trigger beats a token-budget rule at matched retention cannot be evaluated on the release. We state what corpora should ship to answer it.