论文多源确认

高频令牌掩码优化直接偏好优化

Masking Frequent Tokens Sharpens Direct Preference Optimization

精选理由

研究人员发现高频令牌会稀释偏好信号,提出ADPO方法通过掩码这些令牌,在不增加计算成本的情况下提升了模型对齐效果。

研究人员发现直接偏好优化(DPO)中存在高频令牌主导累积序列分数的问题。在Anthropic HH-RLHF数据集上,仅69个令牌类型占所有响应令牌的55.1%和成对共享令牌质量的85.9%。为解决此问题,他们提出了各向异性DPO(ADPO)和频率硬DPO方法。在Qwen-2.5-7B-Instruct和Llama-3-8B-Instruct模型上,该方法在AlpacaEval、MT-Bench和Arena-Hard基准测试中均优于标准DPO。

原文 · arXiv: Anthropic

Masking Frequent Tokens Sharpens Direct Preference Optimization

Direct Preference Optimization (DPO) aligns language models by optimizing over sequence-level sums of token-wise implicit reward differences. However, we identify a pervasive pathology in this formulation: a disproportionately small subset of high-frequency token types dominates cumulative sequence scores while appearing symmetrically across both preferred and dispreferred responses. Specifically, under canonical Qwen tokenization on Anthropic HH-RLHF, merely 69 token types account for $55.1\%$ of all response tokens and $85.9\%$ of within-pair shared token mass, exhibiting substantially lower preference-side specificity than the remaining vocabulary. This symmetric ubiquity induces gradient entanglement and dilutes the discriminative preference signal propagated through the objective. To resolve this issue, we introduce \emph{Anisotropic DPO} (\textsf{ADPO}) and its canonical realization, \emph{Frequency-Hard DPO}. Using a fixed, label-agnostic vocabulary mask, our method zeroes the implicit reward contribution of high-frequency response tokens while assigning unit weight to informative positions, thereby suppressing gradient interference without modifying preference pairs, discarding context, or introducing learned parameters. Here, \emph{anisotropy} designates non-uniform token-level objective weighting rather than representational geometry. Extensive empirical evaluations on AlpacaEval, MT-Bench, and Arena-Hard demonstrate that Frequency-Hard DPO consistently outperforms standard DPO across Qwen-2.5-7B-Instruct and Llama-3-8B-Instruct, establishing that selectively masking shared high-frequency tokens offers an effective, zero-overhead mechanism for robust preference alignment.