Training LLM agents with reinforcement learning (RL) is bottlenecked by environments, which must provide verifiable rewards, support long-horizon interaction, and scale cheaply. Existing approaches rely on costly human-curated data or on LLM-generated environments that risk hallucinations and benchmark contamination. We show that LLMs can instead be trained into capable search agents using synthetic environments generated entirely by rules, whose generation requires no LLM and has zero marginal cost. We build PhantomEnvironments, multi-turn RL environments from fictional worlds, where agents must search a corpus of templated articles to answer multi-hop questions. Despite sharing no facts with the real world, these strikingly simple environments yield agents that transfer to real-world multi-hop search benchmarks, often outperforming real-world training data on newer benchmarks. Trained agents generalize to unseen fictional universes, and Qwen models learn to scale their search budget roughly linearly with question difficulty, suggesting emergent search scaling from environment interaction alone. Ablating environment complexity reveals that hop count drives transfer more than constraints or comparisons: even the simplest rule-generated environments are a surprisingly effective, free resource for training generalizable LLM agents.
Figures & tables
Figure 1: LLM agents trained with RL in PhantomEnvironments transfer to real-world search.
Data source
LLM synthesizes
Verifiability
Search-R1 & followups ( Jin et al., 2025b )
Human-curated
—
Human
ZeroSearch ( Sun et al., 2025 )
LLM-generated
Search engine responses
LLM-judged
ASearcher ( Gao et al., 2025 )
LLM-generated
QA pairs + trajectories
LLM-judged
WebDancer ( Wu et al., 2026 )
LLM-generated
Browsing trajectories
LLM + rejection
WebSailor-V2 ( Li et al., 2026 )
LLM-generated
QA pairs on Wiki seeds
LLM + rejection
SWiRL ( Goldie et al., 2025 )
LLM-generated
Multi-step trajectories
LLM-judged
Table 1: Data and environments for agent training. Prior work uses real or LLM-generated data; ours is the first rule-generated, LLM-free pipeline, with zero-marginal-cost and exact verifiability.
Model
HotpotQA
2Wiki
MuSiQue
Synth-RM
Synth-SM
FRAMES
Qwen2.5-3B
40.1±1.9
30.3±1.9
21.9±1.6
17.7±1.5
9.9±1.2
12.6±1.0
+ PhantomEnvs
60.9±1.5
56.4±1.5
39.4±2.1
34.6±1.8
34.2±1.4
27.4±1.0
Qwen2.5-7B
44.2±2.0
29.1±1.9
24.5±1.8
23.6±1.7
15.6±1.5
16.2±1.1
+ PhantomEnvs
64.1±2.3
66.7±2.2
44.7±2.8
39.1±1.8
40.6±1.8
35.4±1.1
Llama-3.2-3B
16.8±1.5
11.5±1.3
9.4±1.1
8.6±1.1
3.8±0.7
6.8±0.8
+ PhantomEnvs
55.8±1.6
37.5±1.5
35.9±1.8
30.8±2.1
27.0±1.3
23.4±1.1
Table 2: F1 scores on real-world agentic search benchmarks after RL fine-tuning in PhantomEnvironments. Synthetic training significantly improves performance of all LLM families and sizes. Improvements are largest on the newer and harder benchmarks—SynthWorlds-RM, SynthWorlds-SM, and FRAMES—Llama-3.2-3B-Instruct improves by up to 7.1× on SynthWorlds-SM. We report mean ± standard error over the test sets and two training seeds.
Figure 2: F1 scores steadily improve as training progresses. We evaluate intermediate checkpoints of PhantomEnvironments training runs, and observe steady performance improvements across the board: a sharp increase initially then a steady growth. While some runs saturate, we generally do not see drastic overfitting or collapse to the rule-generated templates of PhantomEnvironments. We report mean ± standard error as the solid line and shaded region.
Corpus
Qwen2.5-7B
+ PhantomEnvs
Benchmark’s own
25.5±0.7
48.4±0.8
All pooled
23.9±0.7
46.5±0.8
Table 3: Average F1 scores of Qwen2.5-7B-Instruct, ablating choice of search corpus. Each benchmark’s questions are answered either against that benchmark’s own corpus, or against a single pooled corpus. Expectedly, agents perform worse against the pooled corpus but the drop is slight, with or without training.
Figure 3: F1 scores and number of search calls as a function of question difficulty (hops). (Left two) We evaluate intermediate training checkpoints on validation questions from the training universe. As fine-tuning progresses, F1 increases across all difficulty levels for both LLMs (darker lines are higher in scores). In the lower panels, we plot the number of search calls as a function of question difficulty. We observe an emergent “search scaling” property in Qwen2.5 models: number of searches scales linearly with environment’s question difficulty. Llama-3.2-3B-Instruct shows partial search scaling, increasing search calls initially and plateauing. (Right two) We evaluate final checkpoints on an unseen universe of the same size as training (Unseen 1K) and a 10× larger universe (Unseen 10K). F1 scores and search scaling behavior in Unseen universes parallel the Training universe, indicating agents acquired generalizable agentic search skill rather than memorizing facts.
Figure 4: (a) Training Qwen2.5-3B-Instruct on real-world NQ+HotpotQA data outperforms PhantomEnvironments on benchmarks in-domain to NQ+HotpotQA (HotpotQA, 2WikiMultihopQA, MuSiQue). (b) Whereas PhantomEnvs is better than real-world training on newer and harder out-of-domain benchmarks (SynthWorlds-RM, SynthWorlds-SM, FRAMES). (c) PhantomEnvs teach agents to be equally performant on SynthWorlds pair: F1 scores are similar for RM and SM versions (KA =0 ), but KA >0 for the base model and NQ+HotpotQA training. See text for details.
Figure 5: Ablation results on synthetic environment complexity axes. We find that Hops is the strongest axis that drives real-world transfer.
Figure 6: F1 score deltas per question type of CofCA benchmark after fine-tuning on environment variants. We plot F1 score deltas on comparison and non-comparison questions: we calculate F1 score difference per question, average deltas for each category, and report the mean ± standard error (paired) in percentage points. Hops +Comparisons environment improves performance on evaluation comparison questions, but linear hops remains dominant otherwise.
Appendix figures & tables10 assets
Supplementary material from the paper’s appendix.
Appendix
Evaluation Setup
HotpotQA
2Wiki
MuSiQue
Synth-RM
Synth-SM
FRAMES
Qwen2.5-3B
Paper setup
40.1±1.9
30.3±1.9
21.9±1.6
17.7±1.5
9.9±1.2
12.6±1.0
Training setup
39.8±2.0
29.3±1.8
18.2±1.5
17.6±1.5
10.7±1.3
11.9±0.9
Trained w/ PhantomEnvs
Paper setup
61.5±1.9
56.0±2.0
41.0±1.9
35.9±1.8
34.8±1.9
27.6±1.4
Training setup
56.9±1.9
54.3±2.1
38.8±2.0
31.5±1.8
31.5±1.8
23.4±1.3
Appendix
Table 4: Ablation results on retriever configuration. In the main text, we evaluate every model with a deliberately stronger setup than it trains with ( Qwen3-Embedding-4B , 20 turns, 32768 context). Here we re-evaluate the same checkpoints under the training setup ( e5-base-v2 , 10 turns, 8192 context). Other index settings remain fixed. The weaker training setup generally lowers every model’s score, regardless of how it was trained. We report results for one training seed—including the Paper setup rows, which therefore differ slightly from Table 2 —with standard errors over the test sets. The other training seed follows the same trend.
Evaluation Setup
HotpotQA
2Wiki
MuSiQue
Synth-RM
Synth-SM
FRAMES
Qwen2.5-7B
Paper setup
44.2±2.0
29.1±1.9
24.5±1.8
23.6±1.7
15.6±1.5
16.2±1.1
Training setup
42.6±2.0
32.4±2.0
22.0±1.7
22.7±1.7
15.0±1.5
15.3±1.1
Trained w/ PhantomEnvs
Paper setup
62.2±1.9
65.0±2.0
47.1±2.0
40.3±1.9
41.8±1.9
34.9±1.4
Training setup
62.7±1.9
58.8±2.1
41.8±2.0
35.6±1.9
33.4±1.9
26.9±1.4
Appendix
Table 5: Ablation results on retriever configuration. In the main text, we evaluate every model with a deliberately stronger setup than it trains with ( Qwen3-Embedding-4B , 20 turns, 32768 context). Here we re-evaluate the same checkpoints under the training setup ( e5-base-v2 , 10 turns, 8192 context). Other index settings remain fixed. The weaker training setup generally lowers every model’s score, regardless of how it was trained. We report results for one training seed—including the Paper setup rows, which therefore differ slightly from Table 2 —with standard errors over the test sets. The other training seed follows the same trend.
Corpus
Qwen2.5-3B
+ PhantomEnvs
Benchmark’s own
22.1±0.6
42.1±0.6
All pooled
21.2±0.6
40.7±0.7
Appendix
Table 6: Average F1 scores of Qwen2.5-3B-Instruct, ablating choice of search corpus.
Corpus
Llama-3.2-3B
+ PhantomEnvs
Benchmark’s own
9.5±0.4
35.1±0.7
All pooled
8.1±0.4
29.9±0.6
Appendix
Table 7: Average F1 scores of Llama-3.2-3B-Instruct, ablating choice of search corpus.
Figure 7: F1 scores and number of search calls as a function of question difficulty (hops). We include fine-grained analysis for Qwen2.5-3B-Instruct and Phi-4-mini-instruct here, see Figure 3 and main text for full details.
Figure 8: Qwen2.5-7B-Instruct performance comparison of real and PhantomEnvironments. See Figure 4 and main text for full details.
Figure 9: Llama-3.2-3B-Instruct performance comparison of real and PhantomEnvironments. In out-of-domain benchmarks, PhantomEnvs training reduces but does not fully close the gap to NQ+HotpotQA training for Llama-3.2-3B-Instruct, unlike Qwen2.5 models. See Figure 4 and main text for full details.
Figure 10: Phi-4-mini-instruct performance comparison of real and PhantomEnvironments. In out-of-domain benchmarks, PhantomEnvs training reduces but does not fully close the gap to NQ+HotpotQA training for Phi-4-mini-instruct, unlike Qwen2.5 models. See Figure 4 and main text for full details.
Figure 11: F1 score deltas per question type of 2WikiMultihopQA benchmark after fine-tuning on Hops, Hops +Comparisons, and Hops +Constraints environments. We plot F1 score deltas on comparison and non-comparison questions: we calculate F1 score difference per question, average deltas for each category, and report the mean ± standard error (paired) in percentage points. Hops +Comparisons environment improves performance on evaluation comparison questions, but linear hops remains dominant otherwise.
Model
PopQA SubEM (%, ↑ )
MMLU Acc (%, ↑ )
Wiki-18 PPL ( ↓ )
Qwen2.5-3B
14.15±0.29
66.37±0.38
13.82±0.01
+ PhantomEnvs
14.39±0.21
66.51±0.28
13.94±0.01
+ NQ+HotpotQA
14.78±0.21
66.38±0.27
13.88±0.01
Qwen2.5-7B
16.51±0.31
74.27±0.35
12.60±0.01
+ PhantomEnvs
16.65±0.23
74.31±0.25
12.79±0.02
+ NQ+HotpotQA
16.44±0.24
74.18±0.25
12.65±0.04
Appendix
Table 8: RL fine-tuning leaves pretrained knowledge and fluency intact. We measure parametric factual recall (PopQA SubEM metric ( Mallen et al., 2023 ) ), general knowledge and reasoning (MMLU accuracy ( Hendrycks et al., 2020 ) ), and language-modeling fluency on Wikipedia-2018 (perplexity, 1M passages randomly sampled from the 21M released by ( Jin et al., 2025b ) ). PopQA does not degrade after fine-tuning (either synthetic or NQ+HotpotQA training), except Phi-4-mini-instruct + PhantomEnvs ( −2.1 pp). MMLU accuracy is unchanged throughout after fine-tuning. Wiki-18 perplexity rises slightly after fine-tuning (1-2% relative, and 7% for Phi-4-mini-instruct + PhantomEnvs). Learning agentic search through RL therefore largely teaches procedural skill, one that can be learned in complement to the LLMs’ memorized parametric knowledge. This is in line with Chen et al. (2025a) , who show that RL fine-tuning can retain parametric knowledge.
Scalable AI agents training relies on interactive environments that faithfully simulate the consequences of agent actions. Manually crafted environments are expensive to build, brittle to extend, and fundamentally limited in diversity. A promising direction is to replace manually crafted environments with LLM-simulated counterparts. However, this paradigm hinges on an unexamined core assumption: LLMs can accurately simulate environmental feedback. In practice, LLM-simulated environments suffer from hallucinations, logical inconsistencies, and silent state drift failures that corrupt agent reward signals and compound the construction costs that the paradigm was designed to eliminate. To address this gap, we propose EnvSimBench with four contributions: 1) We provide the first formal definition and operationalization of Environment Simulation Ability (EnvSim Ability) as a quantifiable research objective. 2) We construct EnvSimBench, a rigorous benchmark covering 400 samples across 167 diverse environments, equipped with verifiable labels and fine-grained difficulty stratification along three axes. 3) Systematic evaluations reveal that all state-of-the-art language models suffer from a universal state change cliff: they achieve near-perfect accuracy on tasks when the environment state remains invariant, yet fail catastrophically when multiple states need simultaneous updates. This finding exposes EnvSim Ability as a critical yet largely unaddressed capability gap. 4) We design a constraint-driven simulation pipeline that substantially reduces hallucination, boosts environment synthesis yield by 6.8%, and cuts costs by over 90%. Overall, EnvSimBench serves as both a diagnostic framework and a practical optimization path for reliable LLM-based environment simulation, establishing a foundation for scalable agent training. Code and data are available at https://github.com/cookieApril/EnvSimBench
Yi Liu, TingFeng Hui, Wei Zhang +4
Beijing University of Posts and Telecommunications · The Hong Kong University of Science and Technology · Chongqing University
A world model predicts environment dynamics based on current observations and actions, serving as a core cognitive mechanism for reasoning and planning. In this work, we investigate how world modeling based on language models can further push the boundaries of general agents. (i) We first focus on building foundation models for agentic environment simulation. We introduce Qwen-AgentWorld-35B-A3B and Qwen-AgentWorld-397B-A17B, the first language world models capable of simulating agentic environments covering 7 domains via long chain-of-thought reasoning. Leveraging more than 10M environment interaction trajectories of 7 domains in real-world environments, we develop Qwen-AgentWorld through a three-stage training pipeline: CPT injects general-purpose world modeling capabilities from the state transition dynamics and augmented professional corpora, SFT activates next-state-prediction reasoning, and RL sharpens simulation fidelity through a tailored framework with hybrid rubric-and-rule rewards. To evaluate language world models, we present AgentWorldBench, a comprehensive benchmark constructed from real-world interactions of 5 frontier models on 9 established benchmarks. Empirical results demonstrate that Qwen-AgentWorld significantly outperforms existing frontier models. (ii) Beyond foundation models, we further investigate two complementary paradigms through which world modeling enhances general agents. First, as a decoupled environment simulator, Qwen-AgentWorld supports scalable and controllable simulation of thousands of real-world environments for agentic RL, yielding gains that surpass real-environment training alone. Second, as a unified agent foundation model, world-model training acts as a highly effective warm-up that improves downstream performance across 7 agentic benchmarks. Code: https://github.com/QwenLM/Qwen-AgentWorld
Reinforcement Learning (RL) has emerged as a powerful training paradigm for LLM-based agents. However, scaling agentic RL for deep research remains constrained by two coupled challenges: hand-crafted synthetic data fails to elicit genuine real-world search capabilities, and real-world search dependency during RL training introduces instability and prohibitive cost, which limits the scalability of Agentic RL. LiteResearcher is a training framework that makes Agentic RL scalable: by constructing a lite virtual world that mirrors real-world search dynamics, we enable a continuously improving training recipe that empowers a tiny search agent to outperform large-scale open-source and commercial models (e.g., Tongyi DeepResearch and Claude-4.5 Sonnet). Specifically, on common benchmarks such as GAIA and Xbench, our LiteResearcher-4B achieves open-source state-of-the-art results of 71.3% and 78.0% respectively, demonstrating that scalable RL training is a key enabler for Deep Research Agents.
Wanli Li, Bince Qu, Bo Pan +5
Zhejiang University · Simplex AI · The Hong Kong Polytechnic University