PhantomEnvironments: 规则生成虚构世界训练LLM智能体
PhantomEnvironments: Training LLM Agents in Fictional Worlds
PhantomEnvironments用规则生成虚构环境训练LLM智能体,零成本且效果惊人,还能迁移到真实世界任务。
PhantomEnvironments研究显示,通过完全由规则生成的合成环境训练LLM智能体,无需LLM生成且零边际成本。该环境基于虚构世界模板文章,要求智能体回答多跳问题。尽管不包含真实世界事实,这些简单环境训练的智能体在真实世界多跳搜索基准上表现优异,甚至在较新基准上超越真实世界训练数据。Qwen模型能根据问题难度线性扩展搜索预算,显示出仅通过环境交互即可涌现的搜索扩展能力。
PhantomEnvironments: Training LLM Agents in Fictional Worlds
Training LLM agents with reinforcement learning (RL) is bottlenecked by environments, which must provide verifiable rewards, support long-horizon interaction, and scale cheaply. Existing approaches rely on costly human-curated data or on LLM-generated environments that risk hallucinations and benchmark contamination. We show that LLMs can instead be trained into capable search agents using synthetic environments generated entirely by rules, whose generation requires no LLM and has zero marginal cost. We build PhantomEnvironments, multi-turn RL environments from fictional worlds, where agents must search a corpus of templated articles to answer multi-hop questions. Despite sharing no facts with the real world, these strikingly simple environments yield agents that transfer to real-world multi-hop search benchmarks, often outperforming real-world training data on newer benchmarks. Trained agents generalize to unseen fictional universes, and Qwen models learn to scale their search budget roughly linearly with question difficulty, suggesting emergent search scaling from environment interaction alone. Ablating environment complexity reveals that hop count drives transfer more than constraints or comparisons: even the simplest rule-generated environments are a surprisingly effective, free resource for training generalizable LLM agents.