Dr. Free: You Don't Need Difficulty Rewards for Self-Evolving Search Agents
Organizations: VCIP, School of Computer Science, Nankai University · Kuaishou Technology · Xiamen University
Abstract
A central limitation of current data-free self-evolution methods for training search agents is their reliance on difficulty-based proposer rewards. These methods reward a proposer for generating questions that challenge a co-evolving solver, using solver difficulty as a proxy for question quality. Yet difficulty alone is insufficient to distinguish questions that require cross-passage evidence from those that are answerable via simpler shortcuts. In addition, measuring difficulty demands repeated solver rollouts for every candidate question, leading to substantial computational costs. In this paper, we introduce \methodname, the first self-evolving search framework that eliminates difficulty-based proposer rewards and directly optimizes for evidence necessity relative to shortcut contexts. Dr. Free samples relational chains from a knowledge graph and pairs them with aligned passages, giving question generation an explicit multi-hop structure. A generated question receives a positive information-gain reward only when the likelihood of the target answer under the complete evidence passages exceeds the maximum likelihood under all evaluated shortcut contexts. Because this signal is computed from teacher-forced likelihoods, it removes the need for pass-rate estimation and reduces proposer training time by over . Experiments on seven open-domain QA benchmarks show that Dr. Free outperforms prior data-free search agents and the supervised baseline, with large improvements on multi-hop QA benchmarks.
Figures & tables
| Methods | NQ | TriviaQA | PopQA | HotpotQA | 2Wiki | MuSiQue | Bamboogle | Avg. |
| Qwen2.5-3B-Instruct | ||||||||
| Few-shot / Supervised | ||||||||
| IRCoT ( Trivedi et al., 2023 ) | 0.111 | 0.312 | 0.200 | 0.164 | 0.171 | 0.067 | 0.240 | 0.181 |
| Search-o1 ( Li et al., 2025 ) | 0.238 | 0.472 | 0.262 | 0.221 | 0.218 | 0.054 | 0.320 | 0.255 |
| SFT | 0.249 | 0.292 | 0.104 | 0.186 | 0.248 | 0.044 | 0.112 | 0.176 |
| R1-Instruct ( Guo et al., 2025 ) | 0.210 | 0.449 | 0.171 | 0.208 | 0.275 | 0.060 | 0.192 | 0.224 |
| NQ | TriviaQA | PopQA | HotpotQA | 2Wiki | MuSiQue | Bamboogle | Avg. | |
| Dr. Free | 0.380 | 0.572 | 0.416 | 0.353 | 0.390 | 0.172 | 0.344 | 0.375 |
| Difficulty Only | 0.362 | 0.568 | 0.405 | 0.324 | 0.357 | 0.121 | 0.264 | 0.343 |
| IG + Difficulty | 0.358 | 0.563 | 0.389 | 0.325 | 0.357 | 0.137 | 0.248 | 0.340 |
| w/o KG | 0.366 | 0.536 | 0.365 | 0.319 | 0.356 | 0.129 | 0.272 | 0.335 |
| w/o Proposer Training | 0.377 | 0.570 | 0.393 | 0.289 | 0.281 | 0.063 | 0.168 | 0.306 |
| NQ | TriviaQA | PopQA | HotpotQA | 2Wiki | MuSiQue | Bamboogle | Avg. | |
| Dr. Free (full ) | 0.380 | 0.572 | 0.416 | 0.353 | 0.390 | 0.172 | 0.344 | 0.375 |
| w/o one-shot | 0.366 | 0.564 | 0.390 | 0.324 | 0.373 | 0.142 | 0.360 | 0.360 |
| w/o closed-book | 0.341 | 0.556 | 0.370 | 0.315 | 0.326 | 0.117 | 0.288 | 0.330 |
| w/o per-hop | 0.321 | 0.535 | 0.355 | 0.290 | 0.311 | 0.114 | 0.320 | 0.321 |
| w/o source-only | 0.326 | 0.527 | 0.356 | 0.295 | 0.303 | 0.104 | 0.304 | 0.316 |
| Hop label | # Pairs | Zero-search | Answer in seed |
| 1 | 13,722 | — | 96.9% |
| 2 | 10,131 | 100.0% | 96.4% |
| 3 | 6,630 | 99.98% | 96.3% |
| 4 | 3,440 | 100.0% | 96.0% |
| 20,201 | 99.995% | 96.4% |
| Method | Closed-book shortcut | Source-only shortcut | One-shot shortcut | Single-component shortcut |
| Dr. Zero | 17% [10.9, 25.5] | 81% [72.2, 87.5] | 65% [55.3, 73.6] | 88% [80.2, 93.0] |
| KG + Difficulty | 48% [38.5, 57.7] | 20% [13.3, 28.9] | 45% [36.0, 55.2] | 51% [41.3, 60.6] |
| Dr. Free | 5% [2.2, 11.2] | 3% [1.0, 8.5] | 2% [0.6, 7.0] | 5% [2.2, 11.2] |
| Dr. Free vs. baseline ( -value) | ||||
| vs. Dr. Zero | ||||
| vs. KG + Difficulty | ||||
Appendix figures & tables14 assets
Supplementary material from the paper’s appendix.
Appendix
| Hyperparameter | Setting |
| Algorithm | HRPO |
| Maximum training steps | 50 |
| Optimizer | AdamW |
| Optimizer momentum | |
| Learning rate | |
| Warmup ratio | 0.03 |
| Setting | NQ | TriviaQA | PopQA | HotpotQA | 2Wiki | MuSiQue | Bamboogle | Avg. |
| Ratio 1:1:1:1 | 0.377 | 0.570 | 0.403 | 0.338 | 0.352 | 0.141 | 0.264 | 0.349 |
| Ratio 4:3:2 | 0.375 | 0.572 | 0.408 | 0.324 | 0.354 | 0.129 | 0.296 | 0.351 |
| Ratio 1:1:1 | 0.380 | 0.572 | 0.416 | 0.353 | 0.390 | 0.172 | 0.344 | 0.375 |
| GeneralQA | Multi-HopQA | |||||||
| Method | NQ | TriviaQA | PopQA | HotpotQA | 2Wiki | MuSiQue | Bamboogle | Avg. |
| Qwen2.5-3B-Instruct | ||||||||
| Dr. Zero | 39.7 / | 57.2 / | 43.1 / | 29.8 / | 29.1 / | 9.1 / | 20.0 / | 32.6 / |
| Dr. Free(Ours) | 38.0 / 47.5 | 57.2 / 65.0 | 41.6 / 45.8 | 35.3 / 46.1 | 39.0 / 45.2 | 17.2 / 25.0 | 34.4 / 47.9 | 37.5 / 46.1 |
| Qwen2.5-7B-Instruct | ||||||||
| Dr. Zero | 40.6 / | 60.8 / | 41.6 / | 36.2 / | 34.7 / | 10.4 / | 36.0 / | 37.2 / |
| NQ | TriviaQA | PopQA | HotpotQA | 2Wiki | MuSiQue | Bamboogle | Avg. | |
| Qwen2.5-3B-Instruct | ||||||||
| Dr. Free Iter 1 | 0.358 | 0.563 | 0.399 | 0.336 | 0.364 | 0.138 | 0.320 | 0.354 |
| Dr. Free Iter 2 | 0.365 | 0.565 | 0.413 | 0.344 | 0.379 | 0.141 | 0.320 | 0.362 |
| Dr. Free Iter 3 | 0.380 | 0.572 | 0.416 | 0.353 | 0.390 | 0.172 | 0.344 | 0.375 |
| Qwen2.5-7B-Instruct | ||||||||
| Dr. Free Iter 1 | 0.382 | 0.616 | 0.404 | 0.376 | 0.345 | 0.165 | 0.432 | 0.389 |
| Category | Representative Relations |
| Authorship & creation | author (P50), creator (P170), composer (P86), director (P57), |
| screenwriter (P58), performer (P175) | |
| Geographic & institutional | country (P17), headquarters location (P159), |
| country of origin (P495), location of formation (P740) | |
| Affiliations & roles | occupation (P106), |
| educated at (P69), employer (P108), position held (P39) |
| Case | Shortcut context | Gain without context | Gain with full set | Effect |
| Lake Ontario Saint Lawrence | Source-only | False positive reject | ||
| Hypatima Macquarie River | Individual passage | False positive reject |
| Case | Hop label | Searches | Generated question | Answer | Evidence in seed passage |
| A1 | 3 | 0 | What is the fundamental group of a topological group G? | Abelian | “The fundamental group of a topological group G is abelian.” |
| A2 | 2 | 0 | At what speed range can the Arriflex D-20 camera operate in RAW format? | 1 to 60 frame/s | “The camera is capable of running at speeds from 1 to 60 frame/s.” |
| A3 | 4 | 0 | On which beach did the 2/17th Infantry Battalion first land during the Allied invasion of Lae? | Red Beach | “The 2/17th Infantry Battalion came ashore on Red Beach behind the 2/15th.” |
| Generated question | Construction chain | Answer | |
| 2-hop generations | |||
| (1) | Which high-speed rail line serves the station that is a terminal for the Ōfunato Line? | Ōfunato Line Ichinoseki Station Tōhoku Shinkansen | Tōhoku Shinkansen |
| (2) | Which family of the payload specialist that flew on STS-41-D? | STS-41-D Charles D. Walker Walker family | Walker family |
| (3) | Which country is the home of the island where the Abruka rear lighthouse is located? | Abruka rear lighthouse Abruka Estonia | Estonia |
| 3-hop generations | |||
| (4) | In which country is the shipwreck discussed in Hugh Edwards’s book located? | Hugh Edwards Islands of Angry Ghosts Batavia Australia | Australia |
| Training source | Similar Q–A pairs | Exact questions |
| Dr. Free | 25 (0.048%) | 0 (0.000%) |
| Search-R1 | 1,108 (2.143%) | 14 (0.027%) |
| Method / Run | NQ | TriviaQA | PopQA | HotpotQA | 2Wiki | MuSiQue | Bamboogle | Avg. |
| Qwen2.5-3B-Instruct | ||||||||
| Dr. Zero (reported) | 0.397 | 0.572 | 0.431 | 0.298 | 0.291 | 0.091 | 0.200 | 0.326 |
| Dr. Free (Run 1) | 0.3765 | 0.5684 | 0.4057 | 0.3545 | 0.4008 | 0.1671 | 0.3440 | 0.3739 |
| Dr. Free (Run 2) | 0.3740 | 0.5777 | 0.4166 | 0.3357 | 0.3713 | 0.1514 | 0.3200 | 0.3638 |
| Dr. Free (Run 3) | 0.3795 | 0.5724 | 0.4159 | 0.3530 | 0.3903 | 0.1721 | 0.3440 | 0.3753 |
| Dr. Free (Mean Std.) | ||||||||
| Condition | Context and judge mode | Shortcut decision rule |
| Closed-book | No evidence is supplied. The judge may answer using its parametric knowledge. | The judge returns answerable=true , and its answer matches the hidden gold. |
| Source-only | Only the seed document is supplied under the evidence-grounded prompt. | The judge returns answerable=true and unique=true , and its answer matches the hidden gold. |
| One-shot | The judge receives the top three passages from one retrieval using the training retriever. Each passage is truncated to 340 whitespace-delimited words. | The judge returns answerable=true and unique=true , and its answer matches the hidden gold. |
| Single-component | Each evidence-chain passage is supplied independently under the evidence-grounded prompt. | A shortcut is recorded if any component independently yields a unique answer matching the hidden gold. |
| Core proposer instructions. Placeholders are instantiated for each sampled knowledge-graph chain. | |||
| Task definition | |||
| You are an expert question writer. You are given a source document and the evidence for every hop of a pre-validated chain. All evidence is already supplied: do not search or emit <tool_call> . Compose ONE question (Q) and its single unambiguous answer (A), grounded in this evidence. Output exactly one assistant turn in the form <think>...</think><question>...</question><answer>...</answer> . | |||
| Dynamic input | |||
| Source document (Hop 0) Ordered entity chain Per-hop evidence {document} {chain_summary} {evidence_block} Treat the passages as the complete construction evidence available to the proposer. | |||
| Question rules | |||
| Q1. | Real question. Begin with a wh-word or auxiliary verb, end with exactly one question mark, and match the wh-word to the answer type. | Q2. | No imperative. Never begin with Identify , Name , List , Describe , Explain , State , Provide , or Give . |