论文:将 Jev 决策模型用作强化学习中的参考策略与探索裁判
Can Jev be Your Q or Policy in Reinforcement Learning?
有意思的尝试:Jev 不生成文本、一次前向就出答案,论文让它当参考策略、探索裁判和回放评分器,MiniGrid 和 Atari 上都跑赢了标准 RL 学习器。
arXiv 论文(编号 2610.11692)研究了 Jev 这一单次前向传播、不生成文本的决策模型能否融入强化学习训练。作者分析后指出 Jev 可以满足 RL 系统对答案的绝大部分要求,唯独不能充当价值函数。论文据此构建了三种角色:参考策略、探索裁判和回放评分器,在 9 个 MiniGrid 任务和 3 个 Atari 游戏上验证,训练效果超过标准 RL 学习器,且 Jev 本身全程保持冻结、未被训练。这是 Jev 首次被纳入 RL 学习过程内部。
Can Jev be Your Q or Policy in Reinforcement Learning?
Foundation models supply reinforcement learning (RL) with priors that mitigate its longstanding weaknesses in sample efficiency and transfer, but their token-by-token generation makes queries sequential and costly. Jev, a recently released decision model, generates nothing and returns calibrated, typed answers in a single forward pass. Existing work studies foundation models in RL either as models to be trained or as generators to be prompted, and Jev belongs to neither category, having so far served only as a black box in single domains. How well such a model decides on its own in RL environments, and how it can improve RL as a component of training, therefore remain unaddressed. To this end, in this paper we first examine the requirements that the objects of an RL system place on the answers they consume, and establish that Jev can fulfill all of them except the cardinal use of a value function. The remaining objects form positions that admit several roles each. We then construct algorithms that employ Jev at three of these positions, as a reference policy, an exploration judge, and a replay rater, to improve sample efficiency, exploration, and learning performance. Across nine MiniGrid tasks and three Atari games, training with Jev outperforms a standard RL learner, including where the learner makes no progress alone, while the model itself remains untrained. To our knowledge, we present the first use of Jev within the RL learning process and establish a frozen decision model as a usable component of RL training, inviting further exploration of how Jev and other advanced decision models can improve RL.