RLCP:用校准剪枝让序列推荐的动作集随会话动态伸缩
Reinforcement Learning with Conformal Action Sets: An Application to Sequential Recommendation
做推荐系统的朋友可以看看这篇:RLCP 让推荐候选集大小随会话动态变化,在 KuaiRand-Pure 和 MovieLens 1M 上多样性最高到基线的 5.21 倍。
论文提出 RLCP(Reinforcement Learning with Calibrated Pruning),通过 critic 分数和一个在线更新的阈值来动态调整序列推荐中保留的动作集大小,替代固定 slate 尺寸。阈值根据二值反馈更新,作者证明了沿自适应轨迹观测到的代理未命中率存在确定性上界。论文还推导出价值损失的精确分解,拆分为过滤损失和选择损失,并给出不要求学习参数收敛的有限会话奖励上界。在 KuaiRand-Pure 和 MovieLens 1M 两个数据集上,19 个配置中的每个都至少有一个 RLCP 变体的目录多样性最高,达到最强基线的 1.11 倍到 5.21 倍。
Reinforcement Learning with Conformal Action Sets: An Application to Sequential Recommendation
Sequential recommenders typically use a fixed slate size even though the number of useful alternatives changes within a session. We propose Reinforcement Learning with Calibrated Pruning (RLCP), which adapts the retained action set using critic scores and an online threshold. The threshold is updated from binary feedback indicating whether the set contains an action in a proxy target. We prove a deterministic bound on the observed proxy miss rate along adaptive trajectories. To quantify the effect of pruning on reward, we derive an exact decomposition of value loss into filtering and selection losses. Under explicit proxy and critic approximation conditions, this decomposition yields a finite session reward bound that also accounts for imperfect selection and set truncation, without requiring the learning parameters to converge. Experiments on KuaiRand-Pure and MovieLens 1M compare two RLCP implementations with four RL baselines. In each of the 19 configurations, at least one RLCP variant achieves the highest catalog diversity, reaching $1.11\times$ to $5.21\times$ that of the strongest baseline, with competitive session depth and no larger retained sets.