← Back to explorer

PhantomEnvironments: Training LLM Agents in Fictional Worlds

Type
paper
Venue
arXiv:2609.40221 (cs.LG), submitted 30 Sep 2026
Year
2026
Source
arxiv
Access
free
Language
English
Added
2026-10-02
Verified
2026-10-02

Summary

Training LLM agents with RL is bottlenecked by environments, which must provide verifiable rewards, support long-horizon interaction, and scale cheaply. Existing approaches use costly human-curated data or LLM-generated environments that risk hallucinations and benchmark contamination. PhantomEnvironments builds multi-turn RL environments from fictional worlds generated entirely by rules — templated articles and multi-hop questions (e.g., 'Who is the mother of Alice's father?') with no LLM in generation, so zero marginal cost, fully verifiable, free of distillation, and immune to contamination. Despite sharing no facts with the real world, these strikingly simple environments teach the generalizable skill of agentic search (decompose a question, retrieve documents, compose knowledge) that transfers to real-world multi-hop search benchmarks, often outperforming real-world training data on newer benchmarks. Trained agents generalize to unseen fictional universes, and Qwen models learn to scale their search budget roughly linearly with question difficulty, suggesting emergent search scaling from environment interaction alone. Ablating environment complexity shows hop count drives transfer more than constraints or comparisons. The announcement headline: a 7B LLM trained with RL in PhantomEnvs performs like an agent 10x its size.

Keywords

PhantomEnvironments · synthetic RL environments · LLM agents · search agents · fictional worlds · PhantomWiki · emergent search scaling

Topics

synthetic RL environments, agent training, search agents, fictional worlds, PhantomWiki, generalization, self-improving agents

Research notes

  • Discovery: @anmolkabra (Anmol Kabra, Cornell PhD, SnorkelAI) X post 2026-10-02 4:40 PM ET (https://x.com/anmolkabra/status/2106122141456613555)
  • Paper: arxiv.org/abs/2609.40221
  • Code: github.com/kilian-group/phantom-envs (announced as 'coming soon' — finishing a hero run on new NVIDIA compute at post time)
  • Milestone in Kabra's PhD; latest in the Phantom* line: PhantomWiki (evaluating reasoning in isolation from memorization), PhantomReasoning (post-training LLMs on synthetic data), PhantomEnvironments (training search agents on rule-generated synthetic RL environments)
  • Won a large NVIDIA GPU compute grant to continue synthetic-data/self-improving-agent work
  • Tour: oral at SnorkelAI Frontier Data Summit (Oct 8), oral at COLM LSEI workshop + poster at LLA workshop (Oct 9)
  • Joint work with @SnorkelAI support; contributors @GongAlbert, @ChaoWan0331, @katielulula, @KilianQW
  • License: CC BY 4.0