内在奖励何时促进探索
When Do Intrinsic Rewards Lead to Exploration?
一篇强化学习探索策略的论文,分析了内在奖励的局限性并提出了改进方案。
该研究提出了一种探索的正式标准,通过反事实信息获取来比较策略。研究者在单一环境中展示了计数、预测误差、赋能和信息增益等目标存在帕累托次优问题。论文解释了这些失败原因,并建立了现有内在奖励成功促进最优探索的条件。
When Do Intrinsic Rewards Lead to Exploration?
Intrinsic rewards are designed to guide exploration in reinforcement learning by assigning value to an agent's experience, for example through prediction error or learning progress. However, maximizing these rewards need not produce the most informative experience available. We propose a formal criterion for exploration that compares policies by the counterfactual information they acquire: how well their histories can substitute for experience under alternative policies. We construct a single, simple environment in which specified count-based, prediction-error, empowerment, and information-gain objectives have maximizing policies that are Pareto-suboptimal at acquiring counterfactual information. We explain these failures and establish conditions under which existing intrinsic rewards successfully encourage optimal exploration. We also construct an objective that assigns a higher value whenever exploration strictly improves under our criterion.