论文

Jaxolotl:LTL多任务RL统一基准套件

Jaxolotl: A Unified High-Performance Benchmark Suite for LTL-Based Multi-Task RL

精选理由

Jaxolotl基准套件让多任务RL研究有了统一标准,训练速度提升220倍,还能对比不同算法的优缺点。

Jaxolotl是一个高性能基准套件,专为基于线性时序逻辑(LTL)的多任务强化学习设计。该套件包含六种代表性算法和四种环境的JAX实现,支持JIT编译训练,速度提升最高达220倍。研究团队通过系统评估发现,具有非短视推理能力的方法在命题数量增加时表现下降,而具有更强扩展性的方法则依赖环境特定假设并受限于短视问题。

原文 · arXiv cs.AI

Jaxolotl: A Unified High-Performance Benchmark Suite for LTL-Based Multi-Task RL

Training agents to follow arbitrary instructions is an important goal of multi-task reinforcement learning (RL). Linear temporal logic (LTL) provides a precise and structured formalism for specifying instructions to agents, and has been successfully adopted for training generalist multi-task policies. However, differences in implementations, task distributions, and evaluation protocols make existing methods difficult to compare, while high computational costs limit the scale and statistical reliability of experiments. We introduce Jaxolotl, a unified high-performance benchmark suite for multi-task LTL-RL to address these concerns. Jaxolotl provides a modular, end-to-end JAX implementation of six representative algorithms and four environments, together with newly curated task suites and a standardised, statistically robust evaluation protocol. By precompiling symbolic task representations into static arrays, Jaxolotl enables fully JIT-compiled training and evaluation, achieving end-to-end speedups of up to $220\times$ and supporting controlled comparisons at substantially greater experimental scale. We use this framework to systematically evaluate existing approaches, revealing complementary strengths and limitations: general methods capable of non-myopic reasoning struggle as the number of propositions grows, while methods with stronger scaling rely on environment-specific assumptions and suffer from myopia.