LMArena出了个AutoEval,拿历史对战数据训练奖励模型,几小时就能出评估结果,比之前快多了,搞模型的可以看看。
Arena AutoEval利用数千万条真实人类对战数据训练奖励模型,能在数小时内完成新模型评估,而传统流程通常需要数周。该方法通过人类提示和模型响应进行评分,保持实时基准特性,不同于LLM-as-judge方式。Arena ML工程师Haoran Guo和Wayne Zhang在视频中讲解了其方法论与应用场景。
Arena AutoEval can quickly evaluate new models, allowing model labs to leverage results faster. It c...
Arena AutoEval can quickly evaluate new models, allowing model labs to leverage results faster. It compresses the evaluation feedback loop from weeks to hours, enabling rapid, human-preference-based insight into a model’s capabilities. For more details, see the full video in the post below. Your browser does not support the video tag. 🔗 View on Twitter Arena.ai @arena Arena AutoEval uses tens of millions of historical Arena battles from humans, to train a reward model that can estimate human preferences at scale, and deliver evaluation results at frontier speed. Unlike LLM-as-judge methods, AutoEval remains a live benchmark: it is derived from existing real-world battles, with the reward model scoring from human prompts and model responses. Arena ML Engineers, Haoran Guo and Wayne Zhang, discuss AutoEval below. Watch the full video to learn more about AutoEval’s methodology and model lab use cases in the post below. Your browser does not support the video tag. 🔗 View on Twitter 🔗 View Quoted Tweet 💬 2 🔄 2 ❤️ 25 👀 6164 📊 4 ⚡