研究人员提出PLC-DPO,能校正噪声偏好数据,在57个测试单元中胜率超DPO 5个百分点。
PLC-DPO方法通过将每对训练信号路由为清洁、翻转或平局情况,优化偏好学习。该方法在57个数据集-模型-基准测试单元中,平均胜率达60.5%,超过DPO方法的55.5%。注入噪声测试和人类分歧分析表明,该方法能稳定区分翻转与弱方向性偏好对。
PLC-DPO: Posterior Label Correction in Noisy and Ambiguous Preference Optimization
Direct Preference Optimization (DPO) simplifies alignment through pairwise comparisons but assumes all observed preferences are reliable. Real data often violates this assumption, leading to reversed, weak, or ambiguous labels that cause harmful policy updates. To address this, we propose Posterior Label Correction DPO (PLC-DPO) to robustly optimize preferences by routing each pair's training signal as a clean, flip, or tie case. The key idea is to use the calibrated policy-reference margin as online evidence to take appropriate correction actions. This reframes noisy preference learning as actively correcting supervision direction and strength rather than merely filtering suspicious examples. Across 57 dataset-model-benchmark cells, PLC-DPO obtains the best mean win rate against DPO (60.5 vs. 55.5 for the next-best method). Injected-noise and tie stress tests, human disagreement analysis, and self-confirmation diagnostics further show that the routing remains stable and distinguishes flipped from weakly directional pairs.