Arena训练多模态奖励模型,文本到图像MMRB2上领先9分
Arena这个多模态奖励模型真强,文本到图像那块在MMRB2上比第二名高9分多,挺值的看的。
Arena在视觉、图像生成和代码领域基于实时偏好训练模态特定奖励模型。文本到图像奖励模型使用超过300万偏好对训练。在公开MMRB2基准上,该模型达到点式奖励模型最先进性能,比次优基线高9分以上。
The same approach extends beyond text. Across vision, image generation, and code arena, we train modality-specific reward models on live Arena preferences. For example, on Text-to-Image, our reward model was trained on more than 3 million preference pairs. On the public MMRB2, it achieves state-of-the-art performance among pointwise reward models, more than 9 points ahead of the next-best baseline. 💬 1 🔄 1 ❤️ 3 👀 1183 📊 2 ⚡