cs.AISep 27, 2026

Dr. Free: You Don't Need Difficulty Rewards for Self-Evolving Search Agents

Authors: Zhipeng Qian, Zihan Liang, Yufei Ma, Jie Ma, Ben Chen, Huangyu Dai, Lingtao Mao, Xinyu Sun, +5 more

Organizations: VCIP, School of Computer Science, Nankai University · Kuaishou Technology · Xiamen University

Abstract

A central limitation of current data-free self-evolution methods for training search agents is their reliance on difficulty-based proposer rewards. These methods reward a proposer for generating questions that challenge a co-evolving solver, using solver difficulty as a proxy for question quality. Yet difficulty alone is insufficient to distinguish questions that require cross-passage evidence from those that are answerable via simpler shortcuts. In addition, measuring difficulty demands repeated solver rollouts for every candidate question, leading to substantial computational costs. In this paper, we introduce \methodname, the first self-evolving search framework that eliminates difficulty-based proposer rewards and directly optimizes for evidence necessity relative to shortcut contexts. Dr. Free samples relational chains from a knowledge graph and pairs them with aligned passages, giving question generation an explicit multi-hop structure. A generated question receives a positive information-gain reward only when the likelihood of the target answer under the complete evidence passages exceeds the maximum likelihood under all evaluated shortcut contexts. Because this signal is computed from teacher-forced likelihoods, it removes the need for pass-rate estimation and reduces proposer training time by over 7×7\times. Experiments on seven open-domain QA benchmarks show that Dr. Free outperforms prior data-free search agents and the supervised baseline, with large improvements on multi-hop QA benchmarks.

Figures & tables

Appendix figures & tables14 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Knowledge-Graph Paths as Intermediate Supervision for Self-Evolving Search Agents

    May 7, 2026Huyu Wu, Jun Liu, Xiaochi Wei +3Self-Evolving AgentsSearch Agents

  2. SearchMaster: Grounded and Regulated Self-Play for Search Agents

    Aug 3, 2026Wentao Tan, Qiong Cao, Jiaqi Wang +1Search AgentsSelf-Play

  3. EVE-Agent: Evidence-Verifiable Self-Evolving Agents

    May 21, 2026Yamato Arai, Yuma IchikawaSelf-Evolving AgentsEvidence-Grounded Questions