论文

P-TTT:用测试时训练改进个性化奖励建模

Using Context Is Not Enough: Test-Time Training for Personalized Reward Modeling

精选理由

一篇讲个性化奖励建模的论文,指出了 ICL 方法学不到偏好关系的坑,提出的 P-TTT 一次前向就能更新用户权重,做 RLHF 的可以看看。

现有 RLHF 流程大多训练单一奖励模型,忽略了用户偏好的个体差异。常见的个性化奖励模型依赖上下文学习(ICL),把用户历史比较作为偏好对放进上下文,但论文指出这种方式无法真正捕捉这些偏好对中的偏好关系。作者提出 Preference-Aligned Test-Time Training(P-TTT),通过序列级更新与应用操作,把偏好关系编码进用户专属的快权重。P-TTT 无需推理时反向传播,在单次前向传播内即可完成快权重更新,实验显示其大幅超越现有最优方法。

原文 · arXiv cs.AI

Using Context Is Not Enough: Test-Time Training for Personalized Reward Modeling

Reinforcement learning from human feedback (RLHF) aligns large language models (LLMs) with human preferences, yet most pipelines learn a single reward model that overlooks individual differences in preferences. Personalized reward models (PRMs) address this by conditioning rewards on user-specific feedback, most commonly through in-context learning (ICL), where a user's historical comparisons are supplied as contextual preference pairs. However, we identify a key limitation of ICL-based PRMs: they fail to capture the preference relations conveyed by contextual pairs. To address this, we propose Preference-Aligned Test-Time Training (P-TTT), which explicitly encodes these relations into user-specific fast weights for personalized reward prediction. P-TTT introduces sequence-level update and apply operations to match the response-level granularity of preference feedback, together with a preference-aligned objective that directly uses pairwise preference relations to guide fast-weight adaptation. Notably, P-TTT is simple to implement and computationally efficient, updating fast weights within a single forward pass without inference-time backpropagation. Extensive experiments show that P-TTT more effectively captures historical preference relations and outperforms state-of-the-art methods by a large margin.