While reinforcement learning (RL) enhances their ability to plan and reason across retrieval steps, we identify a critical failure mode in this setting: Tool-Call Hacking. Unlike execution-based tools (e.g., code or math), whose effects are directly observable, the weak observability of causal dependencies between retrieved evidence and reasoning under format- and outcome-level supervision enables agents to maximize surface-level reward signals without genuinely grounding their reasoning in the returned evidence. This leads to distinctive pathologies, including mode collapse via tool overuse and hallucinated tool usage where tool calls are largely decorative. To address this issue, we propose Proof-of-Use (PoU), an evidence grounded RL framework that explicitly optimizes the causal dependency from retrieval to reasoning and final answers. PoU re-fomulate a fine-grained stepwise interaction protocol in which agents must auditably cite normalized evidence identifiers. We operationalize this via a multi-objective reward design consisting of: (1) two progressive process rewards that constrain citation validity at intermediate steps; (2) a global Answer--Support Alignment reward that enforces consistency between final answers and retrieved evidence; and (3) a curriculum-style adaptive reward mixing mechanism that smoothly transitions agents from dense process supervision to sparse outcome-based objectives. Extensive experiments show the strong performance of PoU and demonstrate the effectiveness in mitigating tool-call hacking. Beyond this, PoU exhibits a notable emergent property: adaptive and robust tool-usage patterns naturally arise under domain and tool shifts, even though PoU does not explicitly optimize for tool adaptation.
Figures & tables
Figure 1 . Tool-distribution entropy over training. At each training step, we compute the entropy of the tool usage distribution as H~=−∑ipilogpi/logK , where pi is the proportion of calls to tool i and K is the number of available tools. DeepResearcher exhibits a decreasing entropy, which indicates a tool-call hacking pattern (progressive collapsing toward a narrow tool selection). In contrast, PoU shows higher variance in early training, reflecting exploration, and converges to a stable intermediate entropy value, indicating a learned and task-coordinated tool usage pattern.
Figure 2 . The Proof-of-Use Framework
Model
Env.
HotpotQA
2Wiki
F1
LM
F1
LM
RAG
Routing RAG
Multi-source
30.8
36.0
14.4
18.2
All-in-one RAG
Multi-source
34.5
40.2
23.2
26.4
Search-o1 *
Local
31.6
40.8
28.6
32.8
Search-o1
Web Search
33.0
42.4
30.9
37.7
Table 1 . In-domain evaluation on HotpotQA and 2Wiki datasets. All results are reported with F1 and LLM-based evaluation (LM) metrics. Bold numbers indicate the best performance in each column. Underlined numbers indicate the second best.
Method
Env.
NQ
TQ
MuSiQue
Bamboogle
PopQA
F1
LM
F1
LM
F1
LM
F1
LM
F1
LM
RAG
Routing RAG
Multi-source
34.8
42.9
51.0
59.6
6.5
8.2
17.0
17.4
35.1
39.9
All-in-one RAG
Multi-source
37.5
48.3
60.2
68.8
9.7
11.8
23.5
25.1
46.9
48.8
Search-o1 *
Local
34.5
57.4
52.6
61.1
16.8
21.3
35.8
38.4
36.9
42.4
Search-o1
Web Search
32.4
55.1
58.9
69.5
14.7
19.7
46.6
53.6
38.3
43.4
Table 2 . Performance comparison on five out-of-domain QA benchmarks: NQ , TQ , MuSiQue , Bamboogle , and PopQA . All results are reported with F1 and LLM-based evaluation (LM) metrics. Bold numbers indicate the best performance in each column. Underlined numbers indicate the second best.
Figure 3 . Ablation study of PoU. Comparison under different reward configurations: default perturbation budget B=1 , extended B=2 , and variants removing Rpt or Rans .
Dataset
Zero-shot
RAG
DeepResearcher+
DeepResearcher
DeepResearcher (w/ NT)
PoU
PoU (w/ NT)
BioASQ
13.5
25.8
37.1
38.9
31.8
45.6
43.9
PQArefEval
18.4
31.5
42.4
42.1
37.6
48.7
48.3
Table 3 . OOT evaluation on biomedical QA datasets ( BioASQ , PQArefEval ). “w/ NT” denotes the inclusion of an additional negative tool.
Figure 4 . Tool-call ratio (%) distributions across models and settings. For visual clarity, values below 6% are not shown. For brevity, DeepR. denotes DeepResearcher .