Training LLM agents with reinforcement learning (RL) is bottlenecked by environments, which must provide verifiable rewards, support long-horizon interaction, and scale cheaply. Existing approaches rely on costly human-curated data or on LLM-generated environments that risk hallucinations and benchmark contamination. We show that LLMs can instead be trained into capable search agents using synthetic environments generated entirely by rules, whose generation requires no LLM and has zero marginal cost. We build PhantomEnvironments, multi-turn RL environments from fictional worlds, where agents must search a corpus of templated articles to answer multi-hop questions. Despite sharing no facts with the real world, these strikingly simple environments yield agents that transfer to real-world multi-hop search benchmarks, often outperforming real-world training data on newer benchmarks. Trained agents generalize to unseen fictional universes, and Qwen models learn to scale their search budget roughly linearly with question difficulty, suggesting emergent search scaling from environment interaction alone. Ablating environment complexity reveals that hop count drives transfer more than constraints or comparisons: even the simplest rule-generated environments are a surprisingly effective, free resource for training generalizable LLM agents.
Figures & tables
Figure 1: LLM agents trained with RL in PhantomEnvironments transfer to real-world search.
Data source
LLM synthesizes
Verifiability
Search-R1 & followups ( Jin et al., 2025b )
Human-curated
—
Human
ZeroSearch ( Sun et al., 2025 )
LLM-generated
Search engine responses
LLM-judged
ASearcher ( Gao et al., 2025 )
LLM-generated
QA pairs + trajectories
LLM-judged
WebDancer ( Wu et al., 2026 )
LLM-generated
Browsing trajectories
LLM + rejection
WebSailor-V2 ( Li et al., 2026 )
LLM-generated
QA pairs on Wiki seeds
LLM + rejection
SWiRL ( Goldie et al., 2025 )
LLM-generated
Multi-step trajectories
LLM-judged
Table 1: Data and environments for agent training. Prior work uses real or LLM-generated data; ours is the first rule-generated, LLM-free pipeline, with zero-marginal-cost and exact verifiability.
Model
HotpotQA
2Wiki
MuSiQue
Synth-RM
Synth-SM
FRAMES
Qwen2.5-3B
40.1±1.9
30.3±1.9
21.9±1.6
17.7±1.5
9.9±1.2
12.6±1.0
+ PhantomEnvs
60.9±1.5
56.4±1.5
39.4±2.1
34.6±1.8
34.2±1.4
27.4±1.0
Qwen2.5-7B
44.2±2.0
29.1±1.9
24.5±1.8
23.6±1.7
15.6±1.5
16.2±1.1
+ PhantomEnvs
64.1±2.3
66.7±2.2
44.7±2.8
39.1±1.8
40.6±1.8
35.4±1.1
Llama-3.2-3B
16.8±1.5
11.5±1.3
9.4±1.1
8.6±1.1
3.8±0.7
6.8±0.8
+ PhantomEnvs
55.8±1.6
37.5±1.5
35.9±1.8
30.8±2.1
27.0±1.3
23.4±1.1
Table 2: F1 scores on real-world agentic search benchmarks after RL fine-tuning in PhantomEnvironments. Synthetic training significantly improves performance of all LLM families and sizes. Improvements are largest on the newer and harder benchmarks—SynthWorlds-RM, SynthWorlds-SM, and FRAMES—Llama-3.2-3B-Instruct improves by up to 7.1× on SynthWorlds-SM. We report mean ± standard error over the test sets and two training seeds.
Figure 2: F1 scores steadily improve as training progresses. We evaluate intermediate checkpoints of PhantomEnvironments training runs, and observe steady performance improvements across the board: a sharp increase initially then a steady growth. While some runs saturate, we generally do not see drastic overfitting or collapse to the rule-generated templates of PhantomEnvironments. We report mean ± standard error as the solid line and shaded region.
Corpus
Qwen2.5-7B
+ PhantomEnvs
Benchmark’s own
25.5±0.7
48.4±0.8
All pooled
23.9±0.7
46.5±0.8
Table 3: Average F1 scores of Qwen2.5-7B-Instruct, ablating choice of search corpus. Each benchmark’s questions are answered either against that benchmark’s own corpus, or against a single pooled corpus. Expectedly, agents perform worse against the pooled corpus but the drop is slight, with or without training.
Figure 3: F1 scores and number of search calls as a function of question difficulty (hops). (Left two) We evaluate intermediate training checkpoints on validation questions from the training universe. As fine-tuning progresses, F1 increases across all difficulty levels for both LLMs (darker lines are higher in scores). In the lower panels, we plot the number of search calls as a function of question difficulty. We observe an emergent “search scaling” property in Qwen2.5 models: number of searches scales linearly with environment’s question difficulty. Llama-3.2-3B-Instruct shows partial search scaling, increasing search calls initially and plateauing. (Right two) We evaluate final checkpoints on an unseen universe of the same size as training (Unseen 1K) and a 10× larger universe (Unseen 10K). F1 scores and search scaling behavior in Unseen universes parallel the Training universe, indicating agents acquired generalizable agentic search skill rather than memorizing facts.
Figure 4: (a) Training Qwen2.5-3B-Instruct on real-world NQ+HotpotQA data outperforms PhantomEnvironments on benchmarks in-domain to NQ+HotpotQA (HotpotQA, 2WikiMultihopQA, MuSiQue). (b) Whereas PhantomEnvs is better than real-world training on newer and harder out-of-domain benchmarks (SynthWorlds-RM, SynthWorlds-SM, FRAMES). (c) PhantomEnvs teach agents to be equally performant on SynthWorlds pair: F1 scores are similar for RM and SM versions (KA =0 ), but KA >0 for the base model and NQ+HotpotQA training. See text for details.
Figure 5: Ablation results on synthetic environment complexity axes. We find that Hops is the strongest axis that drives real-world transfer.
Figure 6: F1 score deltas per question type of CofCA benchmark after fine-tuning on environment variants. We plot F1 score deltas on comparison and non-comparison questions: we calculate F1 score difference per question, average deltas for each category, and report the mean ± standard error (paired) in percentage points. Hops +Comparisons environment improves performance on evaluation comparison questions, but linear hops remains dominant otherwise.
Appendix figures & tables10 assets
Supplementary material from the paper’s appendix.
Appendix
Evaluation Setup
HotpotQA
2Wiki
MuSiQue
Synth-RM
Synth-SM
FRAMES
Qwen2.5-3B
Paper setup
40.1±1.9
30.3±1.9
21.9±1.6
17.7±1.5
9.9±1.2
12.6±1.0
Training setup
39.8±2.0
29.3±1.8
18.2±1.5
17.6±1.5
10.7±1.3
11.9±0.9
Trained w/ PhantomEnvs
Paper setup
61.5±1.9
56.0±2.0
41.0±1.9
35.9±1.8
34.8±1.9
27.6±1.4
Training setup
56.9±1.9
54.3±2.1
38.8±2.0
31.5±1.8
31.5±1.8
23.4±1.3
Appendix
Table 4: Ablation results on retriever configuration. In the main text, we evaluate every model with a deliberately stronger setup than it trains with ( Qwen3-Embedding-4B , 20 turns, 32768 context). Here we re-evaluate the same checkpoints under the training setup ( e5-base-v2 , 10 turns, 8192 context). Other index settings remain fixed. The weaker training setup generally lowers every model’s score, regardless of how it was trained. We report results for one training seed—including the Paper setup rows, which therefore differ slightly from Table 2 —with standard errors over the test sets. The other training seed follows the same trend.
Evaluation Setup
HotpotQA
2Wiki
MuSiQue
Synth-RM
Synth-SM
FRAMES
Qwen2.5-7B
Paper setup
44.2±2.0
29.1±1.9
24.5±1.8
23.6±1.7
15.6±1.5
16.2±1.1
Training setup
42.6±2.0
32.4±2.0
22.0±1.7
22.7±1.7
15.0±1.5
15.3±1.1
Trained w/ PhantomEnvs
Paper setup
62.2±1.9
65.0±2.0
47.1±2.0
40.3±1.9
41.8±1.9
34.9±1.4
Training setup
62.7±1.9
58.8±2.1
41.8±2.0
35.6±1.9
33.4±1.9
26.9±1.4
Appendix
Table 5: Ablation results on retriever configuration. In the main text, we evaluate every model with a deliberately stronger setup than it trains with ( Qwen3-Embedding-4B , 20 turns, 32768 context). Here we re-evaluate the same checkpoints under the training setup ( e5-base-v2 , 10 turns, 8192 context). Other index settings remain fixed. The weaker training setup generally lowers every model’s score, regardless of how it was trained. We report results for one training seed—including the Paper setup rows, which therefore differ slightly from Table 2 —with standard errors over the test sets. The other training seed follows the same trend.
Corpus
Qwen2.5-3B
+ PhantomEnvs
Benchmark’s own
22.1±0.6
42.1±0.6
All pooled
21.2±0.6
40.7±0.7
Appendix
Table 6: Average F1 scores of Qwen2.5-3B-Instruct, ablating choice of search corpus.
Corpus
Llama-3.2-3B
+ PhantomEnvs
Benchmark’s own
9.5±0.4
35.1±0.7
All pooled
8.1±0.4
29.9±0.6
Appendix
Table 7: Average F1 scores of Llama-3.2-3B-Instruct, ablating choice of search corpus.
Figure 7: F1 scores and number of search calls as a function of question difficulty (hops). We include fine-grained analysis for Qwen2.5-3B-Instruct and Phi-4-mini-instruct here, see Figure 3 and main text for full details.
Figure 8: Qwen2.5-7B-Instruct performance comparison of real and PhantomEnvironments. See Figure 4 and main text for full details.
Figure 9: Llama-3.2-3B-Instruct performance comparison of real and PhantomEnvironments. In out-of-domain benchmarks, PhantomEnvs training reduces but does not fully close the gap to NQ+HotpotQA training for Llama-3.2-3B-Instruct, unlike Qwen2.5 models. See Figure 4 and main text for full details.
Figure 10: Phi-4-mini-instruct performance comparison of real and PhantomEnvironments. In out-of-domain benchmarks, PhantomEnvs training reduces but does not fully close the gap to NQ+HotpotQA training for Phi-4-mini-instruct, unlike Qwen2.5 models. See Figure 4 and main text for full details.
Figure 11: F1 score deltas per question type of 2WikiMultihopQA benchmark after fine-tuning on Hops, Hops +Comparisons, and Hops +Constraints environments. We plot F1 score deltas on comparison and non-comparison questions: we calculate F1 score difference per question, average deltas for each category, and report the mean ± standard error (paired) in percentage points. Hops +Comparisons environment improves performance on evaluation comparison questions, but linear hops remains dominant otherwise.
Model
PopQA SubEM (%, ↑ )
MMLU Acc (%, ↑ )
Wiki-18 PPL ( ↓ )
Qwen2.5-3B
14.15±0.29
66.37±0.38
13.82±0.01
+ PhantomEnvs
14.39±0.21
66.51±0.28
13.94±0.01
+ NQ+HotpotQA
14.78±0.21
66.38±0.27
13.88±0.01
Qwen2.5-7B
16.51±0.31
74.27±0.35
12.60±0.01
+ PhantomEnvs
16.65±0.23
74.31±0.25
12.79±0.02
+ NQ+HotpotQA
16.44±0.24
74.18±0.25
12.65±0.04
Appendix
Table 8: RL fine-tuning leaves pretrained knowledge and fluency intact. We measure parametric factual recall (PopQA SubEM metric ( Mallen et al., 2023 ) ), general knowledge and reasoning (MMLU accuracy ( Hendrycks et al., 2020 ) ), and language-modeling fluency on Wikipedia-2018 (perplexity, 1M passages randomly sampled from the 21M released by ( Jin et al., 2025b ) ). PopQA does not degrade after fine-tuning (either synthetic or NQ+HotpotQA training), except Phi-4-mini-instruct + PhantomEnvs ( −2.1 pp). MMLU accuracy is unchanged throughout after fine-tuning. Wiki-18 perplexity rises slightly after fine-tuning (1-2% relative, and 7% for Phi-4-mini-instruct + PhantomEnvs). Learning agentic search through RL therefore largely teaches procedural skill, one that can be learned in complement to the LLMs’ memorized parametric knowledge. This is in line with Chen et al. (2025a) , who show that RL fine-tuning can retain parametric knowledge.