CLIFT:用保形自验证训练开源网页智能体并替代外部裁判
CLIFT: Conformal Self-Verification for Web Agent Training and Test-Time Scaling
开源网页智能体训练时奖励信号太稀疏?这篇 CLIFT 用保形自验证把裁判反馈变成可复用信号,测试时还能不调 API 就做轨迹筛选,WebArena Infinity 上开源里最強。
CLIFT 是一种面向网页智能体的训练与测试时扩展方法,核心是保形自验证。训练时,智能体对自己 rollout 回答自然语言验证问题,Compositional Conformal Certifier 只保留与训练期裁判一致的问题信号,并按极性感知提升度赋予权重,混入每步奖励。测试时冻结已认证的问题库,用于 Conformal Trajectory Selection(CTS),通过保守多数投票决定是否换用更好的轨迹,全程无需调用外部裁判。在 WebArena Infinity 上,CLIFT 在开源网页智能体中达到最先进水平;在 VisualWebArena 上,用开源模型训练的问题库可迁移到 GPT-5.5 并在标准评测下达到最先进;在 Online Mind2Web 上零样本提升真实网页智能体表现。
CLIFT: Conformal Self-Verification for Web Agent Training and Test-Time Scaling
Open-source web agents are now strong enough to execute realistic browser tasks, but training them with reinforcement learning still depends on weak supervision: binary task success is too sparse for credit assignment, while frontier-language-model judges are too expensive to call at every step and cannot be assumed available at deployment. We introduce CLIFT, a training and test-time scaling method built around conformal self-verification. During training, the agent answers natural-language verification questions about its own rollouts; a Compositional Conformal Certifier keeps only question signals whose URL-conditional evidence agrees with a training-time judge, assigns signed trust weights through polarity-aware lift, and blends the resulting verifier score into per-step rewards in a way that never subtracts from the judge baseline. At test time, the same certified bank is frozen and reused as structured evidence for Conformal Trajectory Selection (CTS): the agent samples a greedy rollout and one or more diverse retries, the self-verifier summarises each URL trace, and a conservative majority-vote rule chooses whether to swap away from the current incumbent without calling any external judge. This single mechanism supports three settings. On WebArena Infinity, CLIFT achieves state-of-the-art performance among open-source web agents. On VisualWebArena, a bank trained with the open model transfers to GPT-5.5 at test time and reaches state-of-the-art performance under the canonical harness. On Online Mind2Web, without training an agent on the benchmark, translating the certified question bank improves a live-web agent in zero-shot evaluation. Together these results position conformal self-verification as a way to turn costly judge feedback into a reusable training signal and a judge-free test-time scaling signal.