PhantomEnvironments: Training LLM Agents in Fictional Worlds
- Type
- paper
- Venue
- arXiv:2609.40221 (cs.LG), submitted 30 Sep 2026
- Year
- 2026
- Source
- arxiv
- Access
- free
- Language
- English
- Added
- 2026-10-02
- Verified
- 2026-10-02
Summary
Training LLM agents with RL is bottlenecked by environments, which must provide verifiable rewards, support long-horizon interaction, and scale cheaply. Existing approaches use costly human-curated data or LLM-generated environments that risk hallucinations and benchmark contamination. PhantomEnvironments builds multi-turn RL environments from fictional worlds generated entirely by rules — templated articles and multi-hop questions (e.g., 'Who is the mother of Alice's father?') with no LLM in generation, so zero marginal cost, fully verifiable, free of distillation, and immune to contamination. Despite sharing no facts with the real world, these strikingly simple environments teach the generalizable skill of agentic search (decompose a question, retrieve documents, compose knowledge) that transfers to real-world multi-hop search benchmarks, often outperforming real-world training data on newer benchmarks. Trained agents generalize to unseen fictional universes, and Qwen models learn to scale their search budget roughly linearly with question difficulty, suggesting emergent search scaling from environment interaction alone. Ablating environment complexity shows hop count drives transfer more than constraints or comparisons. The announcement headline: a 7B LLM trained with RL in PhantomEnvs performs like an agent 10x its size.
Keywords
PhantomEnvironments · synthetic RL environments · LLM agents · search agents · fictional worlds · PhantomWiki · emergent search scaling
Topics
synthetic RL environments, agent training, search agents, fictional worlds, PhantomWiki, generalization, self-improving agents
Research notes
- Discovery: @anmolkabra (Anmol Kabra, Cornell PhD, SnorkelAI) X post 2026-10-02 4:40 PM ET (https://x.com/anmolkabra/status/2106122141456613555)
- Paper: arxiv.org/abs/2609.40221
- Code: github.com/kilian-group/phantom-envs (announced as 'coming soon' — finishing a hero run on new NVIDIA compute at post time)
- Milestone in Kabra's PhD; latest in the Phantom* line: PhantomWiki (evaluating reasoning in isolation from memorization), PhantomReasoning (post-training LLMs on synthetic data), PhantomEnvironments (training search agents on rule-generated synthetic RL environments)
- Won a large NVIDIA GPU compute grant to continue synthetic-data/self-improving-agent work
- Tour: oral at SnorkelAI Frontier Data Summit (Oct 8), oral at COLM LSEI workshop + poster at LLA workshop (Oct 9)
- Joint work with @SnorkelAI support; contributors @GongAlbert, @ChaoWan0331, @katielulula, @KilianQW
- License: CC BY 4.0