Arena 用数千万人类真实对战数据训了个奖励模型,评估速度飞快,还比 LLM-as-judge 更贴近真实偏好。
Arena AutoEval 使用数千万条历史人类对战数据训练奖励模型,以大规模估算人类偏好并快速输出评估结果。与 LLM-as-judge 方法不同,AutoEval 是实时基准,来源于真实对战数据,由奖励模型对用户提示和模型回复打分。该方法由 Arena 机器学习工程师 Haoran Guo 和 Wayne Zhang 介绍。
Arena AutoEval uses tens of millions of historical Arena battles from humans, to train a reward mode...
Arena AutoEval uses tens of millions of historical Arena battles from humans, to train a reward model that can estimate human preferences at scale, and deliver evaluation results at frontier speed. Unlike LLM-as-judge methods, AutoEval remains a live benchmark: it is derived from existing real-world battles, with the reward model scoring from human prompts and model responses. Arena ML Engineers, Haoran Guo and Wayne Zhang, discuss AutoEval below. Watch the full video to learn more about AutoEval’s methodology and model lab use cases in the post below. Your browser does not support the video tag. 🔗 View on Twitter 💬 1 🔄 1 ❤️ 4 👀 2158 📊 2 ⚡