论文精选

PAIR 方法定位长时程智能体中上下文压缩的具体失效点

精选理由

做智能体评测的值得看:PAIR 教你怎么定位压缩到底丢了什么信息,还提醒你压缩先伤稳定性再伤成功率,自己跑 eval 时要多次对比。

这项工作发现上下文压缩会在少数几个具体环节损害长时程智能体的表现,而不是均匀地降低成功率。研究者提出 PAIR,通过让智能体从同一状态分别在有压缩和无压缩下重放,来隔离压缩的真实影响。典型的压缩平均多增加约 5 步操作,比如在 Venmo 任务中摘要丢失了“仅限同事”的过滤条件,把全部 36 笔付款的总额当成了答案。PAIR 诊断出压缩丢弃了哪些信息后,改写压缩提示词的对应段落。在 AppWorld、OfficeBench 和 tau-Bench Retail 上,PAIR 是各压缩方法中任务完成度最稳定的,接近无压缩运行的水平。

原文 · DAIR.AI

Context compression is a huge bottleneck for long-running agents.

This work finds that context compression hurts long-horizon agents at a few specific points.

They propose PAIR, which replays the agent from the same state with and without a given compression, instead of comparing whole runs that differ in many random ways.

A typical compression adds a few extra steps. The large drops in success come from a small number of compression events.

The harmful compressions drop task conditions the agent hasn't resolved yet. In one Venmo task, the summary dropped the "only from coworkers" filter and reported the total of all 36 payments as the answer.

Other compressions reduce API specs the agent already read to vague prose, so the agent reopens the docs and logs in again, which adds about five steps.

PAIR then diagnoses what information those compressions dropped and rewrites the matching sections of the compression prompt. The agent, compressor model, and tools stay fixed.

On AppWorld, OfficeBench, and tau-Bench Retail, it gives the most consistent task completion of any compressed method and comes close to running with no compression at all.

Compression also lowers run-to-run reliability before it makes tasks unsolvable, so check consistency across repeated runs in your own evals.

Paper: https://t.co/rwviRRq5A1

Chat with Paper: https://t.co/abta8Ebmix