Arena这个多模态奖励模型真强,文本到图像那块在MMRB2上比第二名高9分多,挺值的看的。
Arena在视觉、图像生成和代码领域基于实时偏好训练模态特定奖励模型。文本到图像奖励模型使用超过300万偏好对训练。在公开MMRB2基准上,该模型达到点式奖励模型最先进性能,比次优基线高9分以上。
The same approach extends beyond text. Across vision, image generation, and code arena, we train mod...
The same approach extends beyond text. Across vision, image generation, and code arena, we train modality-specific reward models on live Arena preferences. For example, on Text-to-Image, our reward model was trained on more than 3 million preference pairs. On the public MMRB2, it achieves state-of-the-art performance among pointwise reward models, more than 9 points ahead of the next-best baseline. 💬 1 🔄 1 ❤️ 3 👀 1183 📊 2 ⚡