斯坦福团队搞了个新框架,不用传统评分,用细粒度验证加logprob分布,直接刷了四个Agent基准的SOTA,还能顺便提升RL效率。
斯坦福AI实验室提出LLM-as-a-Verifier框架,通过细粒度评分(1-20级代替1-5级)、logprob分布期望、重复评估和标准分解等方法,提升验证扩展效果。在Terminal-Bench V2、SWE-Bench Verified、RoboRewardBench和MedAgentBench四个Agent基准上取得SOTA。细粒度信号还可用于测试时扩展、强化学习和Agent监控,提高样本效率。
Verification has emerged as a new scaling axis 🚀 LLM-as-a-Verifier shows that scaling verification...
Verification has emerged as a new scaling axis 🚀 LLM-as-a-Verifier shows that scaling verification can push performance to SOTA across Terminal-Bench V2, SWE-Bench Verified, RoboRewardBench, and MedAgentBench. Its fine-grained feedback can also serve as a proxy for estimating task progress and improve RL sample efficiency! Congrats to @jackyk02 @shululi256 @pranav_atreya @liu_yuejiang @jyx_su @chelseabfinn @drmapavone @istoica05 @Azaliamirh Jacky Kwok @jackyk02 How can we extract richer signals from AI Feedback? Introducing LLM-as-a-Verifier✨— a simple verification scaling framework that achieves SOTA on agentic benchmarks 🚀 The key idea: - Use fine-grained scoring granularity (e.g., 1-20 instead of the standard 1-5 scale) - Take the expectation over the full logprob distribution of score tokens - Scale repeated evaluation and criteria decomposition You can use these fine-grained signals for more effective test-time scaling, RL, and agent monitoring! It achieves SOTA across Terminal-Bench V2, SWE-Bench Verified, RoboRewardBench, and MedAgentBench 👑 Advised by @Azaliamirh @istoica05 @drmapavone @chelseabfinn 🧵👇 🔗 View Quoted Tweet 💬 2 🔄 6 ❤️ 41 👀 7155 📊 9 ⚡