A central limitation of current data-free self-evolution methods for training search agents is their reliance on difficulty-based proposer rewards. These methods reward a proposer for generating questions that challenge a co-evolving solver, using solver difficulty as a proxy for question quality. Yet difficulty alone is insufficient to distinguish questions that require cross-passage evidence from those that are answerable via simpler shortcuts. In addition, measuring difficulty demands repeated solver rollouts for every candidate question, leading to substantial computational costs. In this paper, we introduce \methodname, the first self-evolving search framework that eliminates difficulty-based proposer rewards and directly optimizes for evidence necessity relative to shortcut contexts. Dr. Free samples relational chains from a knowledge graph and pairs them with aligned passages, giving question generation an explicit multi-hop structure. A generated question receives a positive information-gain reward only when the likelihood of the target answer under the complete evidence passages exceeds the maximum likelihood under all evaluated shortcut contexts. Because this signal is computed from teacher-forced likelihoods, it removes the need for pass-rate estimation and reduces proposer training time by over 7×. Experiments on seven open-domain QA benchmarks show that Dr. Free outperforms prior data-free search agents and the supervised baseline, with large improvements on multi-hop QA benchmarks.
Figures & tables
Figure 1: Dr. Free outperforms Dr. Zero across five backbones while training faster.
Figure 2: Comparison between Dr. Zero-like methods and Dr. Free. The key differences lie in question construction and proposer reward design. Dr. Zero-like methods construct questions through proposer-side search and optimize the proposer with rollout-based difficulty rewards, which may permit zero-search generation with target answers exposed in seed passages. In contrast, Dr. Free constructs questions from verified knowledge-graph chains with aligned passages and replaces difficulty-based rewards with a shortcut-aware information-gain objective that directly promotes dependence on cross-passage evidence.
Figure 3: Overview of Dr. Free. We first sample pre-validated relational chains from a knowledge graph and align each hop with supporting evidence, providing an explicit multi-hop structure for question construction. Given the validated chain and its aligned evidences, the proposer generates a question without performing online multi-turn search. The shortcut-aware information-gain reward then evaluates the generated question by comparing the likelihood of the target answer under the full evidence context against that under the strongest evaluated shortcut context, directly rewarding questions that require cross-passage evidence.
Methods
NQ
TriviaQA
PopQA
HotpotQA
2Wiki
MuSiQue
Bamboogle
Avg.
Qwen2.5-3B-Instruct
Few-shot / Supervised
IRCoT ( Trivedi et al., 2023 )
0.111
0.312
0.200
0.164
0.171
0.067
0.240
0.181
Search-o1 ( Li et al., 2025 )
0.238
0.472
0.262
0.221
0.218
0.054
0.320
0.255
SFT
0.249
0.292
0.104
0.186
0.248
0.044
0.112
0.176
R1-Instruct ( Guo et al., 2025 )
0.210
0.449
0.171
0.208
0.275
0.060
0.192
0.224
Table 1: Performance on seven open-domain QA benchmarks measured by Exact Match (EM). For each model size, methods are grouped as few-shot/supervised or data-free. Baseline results are taken from the original papers under their reported setups.
NQ
TriviaQA
PopQA
HotpotQA
2Wiki
MuSiQue
Bamboogle
Avg.
Dr. Free
0.380
0.572
0.416
0.353
0.390
0.172
0.344
0.375
Difficulty Only
0.362
0.568
0.405
0.324
0.357
0.121
0.264
0.343
IG + Difficulty
0.358
0.563
0.389
0.325
0.357
0.137
0.248
0.340
w/o KG
0.366
0.536
0.365
0.319
0.356
0.129
0.272
0.335
w/o Proposer Training
0.377
0.570
0.393
0.289
0.281
0.063
0.168
0.306
Table 2: Ablation of Dr. Free with Qwen2.5-3B-Instruct. Difficulty Only replaces the IG reward with the pass-rate difficulty reward; IG + Difficulty combines both rewards; w/o KG restores Dr. Zero-style proposer search; and w/o Proposer Training uses the base proposer to generate questions from KG chains without reward-based optimization. Best results are shown in bold.
NQ
TriviaQA
PopQA
HotpotQA
2Wiki
MuSiQue
Bamboogle
Avg.
Dr. Free (full B )
0.380
0.572
0.416
0.353
0.390
0.172
0.344
0.375
w/o one-shot
0.366
0.564
0.390
0.324
0.373
0.142
0.360
0.360
w/o closed-book
0.341
0.556
0.370
0.315
0.326
0.117
0.288
0.330
w/o per-hop
0.321
0.535
0.355
0.290
0.311
0.114
0.320
0.321
w/o source-only
0.326
0.527
0.356
0.295
0.303
0.104
0.304
0.316
Table 3: Leave-one-shortcut-type-out ablation of the shortcut set B . Each row removes one type of shortcut context from the full set. All four shortcut types contribute positively on average, with the largest degradation occurring when the source-only hop passage contexts are removed.
Hop label
# Pairs
Zero-search
Answer in seed
1
13,722
—
96.9%
2
10,131
100.0%
96.4%
3
6,630
99.98%
96.3%
4
3,440
100.0%
96.0%
≥2
20,201
99.995%
96.4%
Table 4: Behavior of the Dr. Zero proposer under two backbone replications. Across both models, most questions labeled as multi-hop are generated without any search, and their answers are frequently recoverable from the seed passage.
Method
Closed-book shortcut ↓
Source-only shortcut ↓
One-shot shortcut ↓
Single-component shortcut ↓
Dr. Zero
17% [10.9, 25.5]
81% [72.2, 87.5]
65% [55.3, 73.6]
88% [80.2, 93.0]
KG + Difficulty
48% [38.5, 57.7]
20% [13.3, 28.9]
45% [36.0, 55.2]
51% [41.3, 60.6]
Dr. Free
5% [2.2, 11.2]
3% [1.0, 8.5]
2% [0.6, 7.0]
5% [2.2, 11.2]
Dr. Free vs. baseline ( p -value)
vs. Dr. Zero
5.7×10−3
3.1×10−33
3.8×10−24
1.4×10−36
vs. KG + Difficulty
7.0×10−13
1.1×10−4
1.5×10−14
3.9×10−14
Table 5: Gold-blind counterfactual audit of four observable reasoning shortcuts. Each method contributes 100 questions with a matched 51/49 two-/three-hop distribution. Brackets report Wilson 95% confidence intervals. The final two rows report one-sided Fisher exact-test p -values for Dr. Free.
Appendix figures & tables14 assets
Supplementary material from the paper’s appendix.
Appendix
Hyperparameter
Setting
Algorithm
HRPO
Maximum training steps
50
Optimizer
AdamW
Optimizer momentum
β1,β2=0.9,0.999
Learning rate
1×10−6
Warmup ratio
0.03
Appendix
Table 6: Representative proposer and solver configurations for Dr. Free.
Setting
NQ
TriviaQA
PopQA
HotpotQA
2Wiki
MuSiQue
Bamboogle
Avg.
Ratio 1:1:1:1
0.377
0.570
0.403
0.338
0.352
0.141
0.264
0.349
Ratio 4:3:2
0.375
0.572
0.408
0.324
0.354
0.129
0.296
0.351
Ratio 1:1:1
0.380
0.572
0.416
0.353
0.390
0.172
0.344
0.375
Appendix
Table 7: Effect of the hop-sampling ratio used to construct proposer training data. All columns report EM under the same training and evaluation protocol. Values are rounded to three decimal places, and bolding is determined using the unrounded values.
GeneralQA
Multi-HopQA
Method
NQ
TriviaQA
PopQA
HotpotQA
2Wiki
MuSiQue
Bamboogle
Avg.
Qwen2.5-3B-Instruct
Dr. Zero
39.7 /
57.2 /
43.1 /
29.8 /
29.1 /
9.1 /
20.0 /
32.6 /
Dr. Free(Ours)
38.0 / 47.5
57.2 / 65.0
41.6 / 45.8
35.3 / 46.1
39.0 / 45.2
17.2 / 25.0
34.4 / 47.9
37.5 / 46.1
Qwen2.5-7B-Instruct
Dr. Zero
40.6 /
60.8 /
41.6 /
36.2 /
34.7 /
10.4 /
36.0 /
37.2 /
Appendix
Table 8: Evaluation with Exact Match (EM) and token-level F1. Each cell reports EM / F1 in percent. Boldface marks the best comparable value within each backbone group. F1 is compared only when both methods report it; blank entries denote results that have not yet been provided.
Figure 4: Anchor-update ablation across two iterations. Panel (a) shows the solver’s raw training rewards (faint) and five-step moving averages; Panel (b) shows held-out validation reward. Iteration 1 is the shared prefix, and the two Iteration-2 strategies branch at cumulative step 50. The isolated excursion at updated-anchor steps 34–36 is omitted from the plot but retained in the accompanying data.
NQ
TriviaQA
PopQA
HotpotQA
2Wiki
MuSiQue
Bamboogle
Avg.
Qwen2.5-3B-Instruct
Dr. Free Iter 1
0.358
0.563
0.399
0.336
0.364
0.138
0.320
0.354
Dr. Free Iter 2
0.365
0.565
0.413
0.344
0.379
0.141
0.320
0.362
Dr. Free Iter 3
0.380
0.572
0.416
0.353
0.390
0.172
0.344
0.375
Qwen2.5-7B-Instruct
Dr. Free Iter 1
0.382
0.616
0.404
0.376
0.345
0.165
0.432
0.389
Appendix
Table 9: Performance of Dr. Free across self-evolution iterations. The best result in each column within each backbone is shown in bold. Both backbones improve across iterations, with the largest gains occurring on multi-hop benchmarks.
Figure 5: Knowledge-graph chain construction pipeline. Random walks on Wikidata5M generate candidate chains, which are filtered for acyclicity, functional or near-functional relations, and unique entity grounding. Retrieval verification in Wiki-18 checks that the source passage mentions the first bridge entity and that subsequent entities are reachable through per-hop retrieval.
Category
Representative Relations
Authorship & creation
author (P50), creator (P170), composer (P86), director (P57),
screenwriter (P58), performer (P175)
Geographic & institutional
country (P17), headquarters location (P159),
country of origin (P495), location of formation (P740)
Affiliations & roles
occupation (P106),
educated at (P69), employer (P108), position held (P39)
Appendix
Table 10: Representative functional and near-functional relation types retained in the whitelist for chain sampling. The whitelist contains approximately 40 relation types across five categories.
Case
Shortcut context
Gain without context
Gain with full set B
Effect
Lake Ontario → Saint Lawrence
Source-only
Δ−src=+1.422
ΔB=−0.953
False positive → reject
Hypatima → Macquarie River
Individual passage
Δ−hop=+9.754
ΔB=−0.117
False positive → reject
Appendix
Table 11: Representative cases showing how the evaluated shortcut contexts prevent false-positive information-gain rewards. Here, Δ−b denotes the gain computed after excluding shortcut context b , whereas ΔB denotes the gain computed using the complete shortcut set defined in equation 5 . Adding the source-only or individual-passage comparison changes an otherwise positive gain into a negative one, revealing that the target answer can already be supported by the corresponding shortcut context.
Case
Hop label
Searches
Generated question
Answer
Evidence in seed passage
A1
3
0
What is the fundamental group of a topological group G?
Abelian
“The fundamental group of a topological group G is abelian.”
A2
2
0
At what speed range can the Arriflex D-20 camera operate in RAW format?
1 to 60 frame/s
“The camera is capable of running at speeds from 1 to 60 frame/s.”
A3
4
0
On which beach did the 2/17th Infantry Battalion first land during the Allied invasion of Lae?
Red Beach
“The 2/17th Infantry Battalion came ashore on Red Beach behind the 2/15th.”
Appendix
Table 12: Representative Dr. Zero proposer failures. Despite their nominal multi-hop labels, all three questions are generated without search and can be answered directly from the seed passage. Questions are reproduced verbatim from the generated data.
Generated question
Construction chain
Answer
2-hop generations
(1)
Which high-speed rail line serves the station that is a terminal for the Ōfunato Line?
Ōfunato Line → Ichinoseki Station → Tōhoku Shinkansen
Tōhoku Shinkansen
(2)
Which family of the payload specialist that flew on STS-41-D?
STS-41-D → Charles D. Walker → Walker family
Walker family
(3)
Which country is the home of the island where the Abruka rear lighthouse is located?
Abruka rear lighthouse → Abruka → Estonia
Estonia
3-hop generations
(4)
In which country is the shipwreck discussed in Hugh Edwards’s book located?
Hugh Edwards → Islands of Angry Ghosts → Batavia → Australia
Australia
Appendix
Table 13: Representative unedited proposer generations and their corresponding construction chains. The chains are included for exposition and are not part of the generated outputs.
Training source
Similar Q–A pairs
Exact questions
Dr. Free
25 (0.048%)
0 (0.000%)
Search-R1
1,108 (2.143%)
14 (0.027%)
Appendix
Table 14: Training–test overlap across 51,713 evaluation questions. A similar Q–A pair has a question Jaccard similarity of at least 0.6 and at least one shared normalized answer alias.
Method / Run
NQ
TriviaQA
PopQA
HotpotQA
2Wiki
MuSiQue
Bamboogle
Avg.
Qwen2.5-3B-Instruct
Dr. Zero (reported)
0.397
0.572
0.431
0.298
0.291
0.091
0.200
0.326
Dr. Free (Run 1)
0.3765
0.5684
0.4057
0.3545
0.4008
0.1671
0.3440
0.3739
Dr. Free (Run 2)
0.3740
0.5777
0.4166
0.3357
0.3713
0.1514
0.3200
0.3638
Dr. Free (Run 3)
0.3795
0.5724
0.4159
0.3530
0.3903
0.1721
0.3440
0.3753
Dr. Free (Mean ± Std.)
0.3767±0.0028
0.5728±0.0047
0.4127±0.0061
0.3477±0.0104
0.3875±0.0150
0.1635±0.0108
0.3360±0.0139
0.3710±0.0063
Appendix
Table 15: Robustness across independent training runs. Dr. Zero results are taken from the single runs reported by Yue et al. (2026) . Dr. Free results are reported for three independent runs, followed by the mean and sample standard deviation. At both model scales, every Dr. Free run exceeds the corresponding average EM reported by Dr. Zero.
Condition
Context and judge mode
Shortcut decision rule
Closed-book
No evidence is supplied. The judge may answer using its parametric knowledge.
The judge returns answerable=true , and its answer matches the hidden gold.
Source-only
Only the seed document is supplied under the evidence-grounded prompt.
The judge returns answerable=true and unique=true , and its answer matches the hidden gold.
One-shot
The judge receives the top three passages from one retrieval using the training retriever. Each passage is truncated to 340 whitespace-delimited words.
The judge returns answerable=true and unique=true , and its answer matches the hidden gold.
Single-component
Each evidence-chain passage is supplied independently under the evidence-grounded prompt.
A shortcut is recorded if any component independently yields a unique answer matching the hidden gold.
Appendix
Table 16: Context construction and decision rules for the gold-blind counterfactual audit.
Core proposer instructions. Placeholders are instantiated for each sampled knowledge-graph chain.
Task definition
You are an expert question writer. You are given a source document and the evidence for every hop of a pre-validated chain. All evidence is already supplied: do not search or emit <tool_call> . Compose ONE question (Q) and its single unambiguous answer (A), grounded in this evidence. Output exactly one assistant turn in the form <think>...</think><question>...</question><answer>...</answer> .
Dynamic input
Source document (Hop 0) Ordered entity chain Per-hop evidence {document} {chain_summary} {evidence_block} Treat the passages as the complete construction evidence available to the proposer.
Question rules
Q1.
Real question. Begin with a wh-word or auxiliary verb, end with exactly one question mark, and match the wh-word to the answer type.
Q2.
No imperative. Never begin with Identify , Name , List , Describe , Explain , State , Provide , or Give .
Appendix
Table 17: Summary of the proposer prompt for question construction. The proposer receives an ordered chain of entities and aligned evidence, and is instructed to generate a question without search while hiding intermediate entities and preserving sequential dependence. Positive examples, counterexamples, and detailed self-checks are omitted for brevity.
Self-evolving search agents reduce reliance on human-written training questions by generating and solving their own search tasks. We build on Search Self-Play (SSP), a representative Proposer and Solver framework in which questions are generated and answered via multi-step search and reasoning. In practice, however, SSP faces two bottlenecks: the Proposer constructs questions from isolated answer entities without relational context, yielding many invalid or unverifiable questions in early self-play training, while the Solver receives only a binary outcome reward that discards useful signal from partially on-track search trajectories. We address both bottlenecks by reusing knowledge-graph paths as construction-derived intermediate supervision for both question construction and reward shaping. First, we ground question construction in LLM-guided knowledge-graph subgraphs, providing relational context for the Proposer. Second, we observe that constructing and solving a multi-hop question can involve overlapping intermediate entities: the factual bridges used to formulate the question may provide approximate waypoints for answering it. Exploiting this overlap, we introduce Waypoint Coverage Reward (WCR), which grants graded partial credit to incorrect Solver trajectories according to their coverage of entities on the construction path, while preserving full reward for correct answers. Across seven QA benchmarks and nine model configurations, our approach improves the average score over standard SSP in all configurations, including notable gains on multi-hop QA tasks. These results suggest that knowledge-graph paths can be reused as lightweight intermediate supervision, providing both relational guidance and process feedback without additional task-specific human annotations or manually labeled process steps.
Training LLM-based search agents requires high-quality search data: tasks that demand genuine multi-hop retrieval and trajectories that use search tools effectively. Existing pipelines often depend on human-written tasks, expert demonstrations, or stronger teacher models. We present SearchMaster, a self-play framework that trains a single LLM from search tasks it generates, solves, and verifies in a local search environment. The key challenge is that self-generated tasks and rollouts can yield misleading signals: pseudo multi-hop questions, success-rate difficulty estimates that ignore search depth, and rollouts with excessive opening but little targeted evidence acquisition. SearchMaster addresses these failure modes with three controls. An Evidence-Chain Generator (ECG) grounds task generation in explicit cross-document evidence chains to reduce pseudo multi-hop questions. A Search-Depth Reward (SDR) scores task difficulty by the search depth of successful rollouts rather than success rate alone, keeping retained tasks search-intensive. An Over-Opening Penalty (OOP) regulates tool use by discouraging excessive document opening, avoiding long but shallow browsing. Verified Proposer and Solver rollouts are then jointly optimized with GRPO. Across six deep-search benchmarks, SearchMaster improves a Qwen3.5-9B backbone from 38.19% to 51.52% average accuracy, with a 30.1-point gain on BrowseComp-Plus. These results show that grounded and regulated self-play can provide effective search-agent training data without human-labeled QA pairs or expert demonstrations. The code is available at https://github.com/WentaoTan/SearchMaster.
Self-evolving agents should not train on examples they cannot justify. Data-free self-evolving search agents offer a scalable route to systems that generate their own questions, answer them, and improve from their own feedback without human annotations. Yet, without verifiable evidence, this loop can reward fluent but unsupported examples, turning the self-generated curriculum into an opaque and potentially unreliable training signal. We argue that evidence verifiability is a prerequisite for trustworthy self-evolution in search agents: each generated instance should include not only an answer but also a source-grounded span whose contribution to that answer can be measured. We introduce EVE-Agent, an Evidence-Verifiable Self-Evolving Agent that operationalizes this principle through a modification to the proposer--solver framework. The proposer generates a question, an answer, and a verbatim evidence span. An evidence verifier then rewards the span according to the marginal accuracy gain when the evidence is provided. This produces a training signal that favors evidence that genuinely helps answer the question, without requiring oracle answers, human labels, or external annotations. EVE-Agent leaves the backbone model, retriever, search tool, and optimization framework unchanged. Experiments show that EVE-Agent substantially improves evidence-grounded correctness over prior self-evolving search agents. The resulting curriculum is not merely self-generated but auditable by construction: each training example carries an inspectable source span that explains why it should be trusted.
Yamato Arai, Yuma Ichikawa
Fujitsu Limited Department of Basic Science · The University of Tokyo · Fujitsu Limited +1