论文71°

OpenBMB推出扩散奖励模型DRM

精选理由

OpenBMB发DRM模型,不再把人类偏好压缩成单一分数,能保留分歧和不确定性,还能提升奖励决策可靠性。

OpenBMB团队发布了扩散奖励模型(DRM),该模型学习完整的奖励分布而非单一分数。DRM能保留人类判断中的分歧、不确定性和多种合理判断。研究显示,DRM可利用分布不确定性识别不稳定的奖励决策,并在RLHF训练中提升下游策略性能。

原文 · OpenBMB

A reward of “3” can mean two completely different things. Everyone thinks a response is mediocre — or half the people love it while the other half hate it. Most Reward Models cannot tell the difference. Introducing Diffusion Reward Models (DRM): instead of collapsing human preference into a single score, DRM learns the full reward distribution, preserving disagreement, uncertainty, and multiple plausible judgments. ✨ Paper:https://t.co/wS1fmHBeRV 🤗 Models: https://t.co/RGTU759A5y 💻 GitHub: https://t.co/KAow6DcoNn

Why it matters: Human disagreement is structured, not just noise. On datasets with repeated annotations, judgments often form separated or polarized patterns. More importantly, as human disagreement increases, DRM’s learned reward distribution becomes increasingly multimodal.

The distribution is useful, not just descriptive. DRM can use distributional uncertainty to identify unstable reward decisions, and distribution-aware ranking improves Best-of-N selection beyond simply taking the mean reward.

Reward Models get their own test-time scaling. Instead of only spending more compute on generating more responses, DRM can keep the response fixed and sample its reward distribution more times. More reward samples give a more reliable estimate — a new scaling axis unavailable to deterministic scalar RMs.

And the gains survive RLHF. When used as the training-time reward, DRM improves downstream policy performance over scalar reward baselines, showing that the benefit is not limited to offline RM benchmarks.

The takeaway: Reward modeling may lose something important when it compresses every human judgment into one number. “Everyone thinks this is average” and “people strongly disagree about this” should not look identical to a Reward Model. DRM makes that difference visible — and usable.