Arena.ai 出了 AutoEval,用真实用户偏好训练奖励模型,几个小时就能评测新模型,比人工快几十倍,结果还跟人工差不多。
Arena.ai 发布了新评测方法 AutoEval,使用基于数百万真实 Arena 用户偏好训练的奖励模型对模型进行排名。该方法与实时人工评测结果高度一致,并将评测速度从数天缩短至数小时。AutoEval 支持文本、视觉、图像和代码 Arena 多种模态。新模型上线后可直接在排行榜显示 AutoEval 预估分数。
这个评测方法有点东西!!! AutoEval:基于数百万真实 Arena 用户偏好的奖励模型对模型进行排名。
这个评测方法有点东西!!! AutoEval:基于数百万真实 Arena 用户偏好的奖励模型对模型进行排名。 Arena.ai @arena Today we’re launching AutoEval: a new evaluation methodology that ranks models using reward models based on millions of real Arena user preferences. Highlights: - High-quality evaluation signals calibrated on real preference data - Strong alignment with live human evaluations - Evaluations that are several orders of magnitude faster (hours instead of days) - Support for Text, Vision, Image, and Code Arena AutoEval enables us to evaluate newly launched models much faster and share results with the community sooner. We’ll now show AutoEval estimated scores for new models directly on the leaderboard. More details in the thread. 🧵 🔗 View Quoted Tweet 💬 1 🔄 0 ❤️ 1 👀 413 📊 1 ⚡