Chai团队搞了个AutoEval奖励模型,预测人类偏好比GPT-4之类还准8-10%,新模型刷榜前先看它评分,省得等人类投票。
AutoEval团队训练了一个奖励模型,专门应对偏好漂移、模型差异细微、平局投票以及长上下文依赖交互等挑战。评估采用时间留出法:用历史投票训练,测试后续发布的模型。早期评分与实时排名的排名相关性超过0.98。该文本奖励模型预测人类偏好的准确率比前沿LLM裁判高8-10%,可作为人类投票积累前的早期排行榜信号。
We train our reward model to handle challenges like preference shifts, increasingly subtle model dif...
We train our reward model to handle challenges like preference shifts, increasingly subtle model differences, ambiguous tie votes, and longer and more context-dependent interactions. To evaluate it fairly, we use a temporal holdout: train on historical votes, and test on models released later. AutoEval’s early scores achieved >0.98 ranking correlation with the live scores that followed. Our text reward model also predicts human preferences 8-10% more accurately than frontier LLM judges. This is what makes AutoEval useful as an early leaderboard signal while human votes are still accumulating. 💬 1 🔄 1 ❤️ 3 👀 751 📊 2 ⚡