论文

Wayfarer:通过拉普拉斯表示学习自动发现选项,攻克 Atari 高难游戏

Mastering Atari 2600 Games with Discovered Options

精选理由

强化学习老难题——自动发现选项,这篇用拉普拉斯表示在 Atari 最难游戏上拿到单流 SOTA,做 RL 的朋友可以细看。

DeepMind 论文提出 Wayfarer,一个在线深度强化学习智能体,通过 Laplacian 表示学习从高维观测中自动发现选项(options)。这些选项同时改善探索、加速信用分配,并能泛化到未见环境。在 Atari 2600 最难的游戏上,Wayfarer 在单流智能体中取得 SOTA,在 Montezuma's Revenge 和 Private Eye 这类需要长视野探索的游戏中提升最大。

原文 · arXiv cs.LG

Mastering Atari 2600 Games with Discovered Options

Temporal abstractions, often instantiated as options, have long been regarded as a mechanism for accelerating credit assignment, facilitating exploration, and enabling generalisation in reinforcement learning (RL). However, developing general option discovery methods that are effective in large-scale, high-dimensional domains remains a fundamental challenge. Existing option discovery methods are either confined to relatively simple domains, depend on handcrafted or quasi-symbolic representations, or offer little improvement over learning without options. We present Wayfarer, a general, domain-agnostic, online deep RL agent that discovers options through Laplacian representation learning from high-dimensional observations and leverages them for control. We show that the resulting options simultaneously improve exploration, accelerate credit assignment, and generalise effectively to unseen settings, enabling substantially faster learning of complex policies. Wayfarer achieves state-of-the-art performance among single-stream agents on the most challenging Atari 2600 games, with the largest gains in games that require long-horizon exploration and strategic behaviour, such as Montezuma's Revenge and Private Eye.