Google 发布 VeriHarness:用智能体验证器替代多数投票选答案
Banger paper from Google. It's standard practice to sample several agent rollouts and trust the ans...
Google 这篇论文有点反直觉:答案一致反而可能是错的。VeriHarness 用同款模型当验证器挑毛病,五个基准上都能涨分,数据也开源了。
Google 论文指出,对多个 agent rollout 取共识答案的做法可能掩盖共同错误,而分歧反而常指向正确答案。VeriHarness 将同一个基础模型变成 agentic verifier:一边在 rollout 分歧时核查 workspace 证据,一边挑战全员一致但可能遗漏需求的结论。在五个 long-horizon 基准上取得最佳选择成绩,配合证据支持的修订,Gemini 3.5 Flash 单 rollout 提升 6.2 分,Claude Opus 4.8 提升 6.4 分。作者同时开源了约 26,000 条 rollout 数据。
Banger paper from Google. It's standard practice to sample several agent rollouts and trust the ans...
Banger paper from Google. It's standard practice to sample several agent rollouts and trust the answers they agree on. This Google paper shows that agreement can hide shared errors, while disagreement often points to the correct alternative. VeriHarness turns the same base model into an agentic verifier with two jobs. One resolves claims where rollouts disagree by checking workspace evidence. The other challenges claims that every rollout agrees on and looks for requirements they all missed. Across five long-horizon benchmarks, it gives the best selection scores among the baselines tested. With evidence-backed revision, it adds 6.2 points over a single rollout with Gemini 3.5 Flash and 6.4 points with Claude Opus 4.8. The authors also release about 26,000 rollouts. Paper: arxiv.org/abs/2610.00972 Chat with Paper: academy.dair.ai/papers/verihar… 💬 8 🔄 1 ❤️ 24 👀 1837 📊 12 ⚡
- arXiv cs.AI10-02 12:42原文