While reinforcement learning (RL) enhances their ability to plan and reason across retrieval steps, we identify a critical failure mode in this setting: Tool-Call Hacking. Unlike execution-based tools (e.g., code or math), whose effects are directly observable, the weak observability of causal dependencies between retrieved evidence and reasoning under format- and outcome-level supervision enables agents to maximize surface-level reward signals without genuinely grounding their reasoning in the returned evidence. This leads to distinctive pathologies, including mode collapse via tool overuse and hallucinated tool usage where tool calls are largely decorative. To address this issue, we propose Proof-of-Use (PoU), an evidence grounded RL framework that explicitly optimizes the causal dependency from retrieval to reasoning and final answers. PoU re-fomulate a fine-grained stepwise interaction protocol in which agents must auditably cite normalized evidence identifiers. We operationalize this via a multi-objective reward design consisting of: (1) two progressive process rewards that constrain citation validity at intermediate steps; (2) a global Answer--Support Alignment reward that enforces consistency between final answers and retrieved evidence; and (3) a curriculum-style adaptive reward mixing mechanism that smoothly transitions agents from dense process supervision to sparse outcome-based objectives. Extensive experiments show the strong performance of PoU and demonstrate the effectiveness in mitigating tool-call hacking. Beyond this, PoU exhibits a notable emergent property: adaptive and robust tool-usage patterns naturally arise under domain and tool shifts, even though PoU does not explicitly optimize for tool adaptation.
Figures & tables
Figure 1 . Tool-distribution entropy over training. At each training step, we compute the entropy of the tool usage distribution as H~=−∑ipilogpi/logK , where pi is the proportion of calls to tool i and K is the number of available tools. DeepResearcher exhibits a decreasing entropy, which indicates a tool-call hacking pattern (progressive collapsing toward a narrow tool selection). In contrast, PoU shows higher variance in early training, reflecting exploration, and converges to a stable intermediate entropy value, indicating a learned and task-coordinated tool usage pattern.
Figure 2 . The Proof-of-Use Framework
Model
Env.
HotpotQA
2Wiki
F1
LM
F1
LM
RAG
Routing RAG
Multi-source
30.8
36.0
14.4
18.2
All-in-one RAG
Multi-source
34.5
40.2
23.2
26.4
Search-o1 *
Local
31.6
40.8
28.6
32.8
Search-o1
Web Search
33.0
42.4
30.9
37.7
Table 1 . In-domain evaluation on HotpotQA and 2Wiki datasets. All results are reported with F1 and LLM-based evaluation (LM) metrics. Bold numbers indicate the best performance in each column. Underlined numbers indicate the second best.
Method
Env.
NQ
TQ
MuSiQue
Bamboogle
PopQA
F1
LM
F1
LM
F1
LM
F1
LM
F1
LM
RAG
Routing RAG
Multi-source
34.8
42.9
51.0
59.6
6.5
8.2
17.0
17.4
35.1
39.9
All-in-one RAG
Multi-source
37.5
48.3
60.2
68.8
9.7
11.8
23.5
25.1
46.9
48.8
Search-o1 *
Local
34.5
57.4
52.6
61.1
16.8
21.3
35.8
38.4
36.9
42.4
Search-o1
Web Search
32.4
55.1
58.9
69.5
14.7
19.7
46.6
53.6
38.3
43.4
Table 2 . Performance comparison on five out-of-domain QA benchmarks: NQ , TQ , MuSiQue , Bamboogle , and PopQA . All results are reported with F1 and LLM-based evaluation (LM) metrics. Bold numbers indicate the best performance in each column. Underlined numbers indicate the second best.
Figure 3 . Ablation study of PoU. Comparison under different reward configurations: default perturbation budget B=1 , extended B=2 , and variants removing Rpt or Rans .
Dataset
Zero-shot
RAG
DeepResearcher+
DeepResearcher
DeepResearcher (w/ NT)
PoU
PoU (w/ NT)
BioASQ
13.5
25.8
37.1
38.9
31.8
45.6
43.9
PQArefEval
18.4
31.5
42.4
42.1
37.6
48.7
48.3
Table 3 . OOT evaluation on biomedical QA datasets ( BioASQ , PQArefEval ). “w/ NT” denotes the inclusion of an additional negative tool.
Figure 4 . Tool-call ratio (%) distributions across models and settings. For visual clarity, values below 6% are not shown. For brevity, DeepR. denotes DeepResearcher .
Large language model (LLM) agents increasingly interleave natural language reasoning with external tools such as web search and code execution. These tool-use policies are often optimized via reinforcement learning (RL), which can amplify spurious correlations in the training data. In this work, we study when and why RL-trained agents learn shortcut tool-selection policies: invoking tools based on superficial prompt cues rather than genuine task requirements. We construct controlled synthetic environments combining factual question answering and mathematical reasoning tasks, and inject cues that are strongly correlated with specific tools during training but causally irrelevant to tool necessity. Across counterfactual evaluations where cues are present but the associated tools are not required, agents exhibit substantial shortcut behavior, with spurious tool invocation rates increasing by up to 39 percent. However, shortcut formation is not universal: across the conditions we test, it arises only when the agent has already learned to use the target tool reliably, suggesting that task competence, rather than dataset imbalance alone, is a key factor in shortcut learning. A swapped-cue analysis further shows that semantic alignment between cues and tools substantially amplifies this effect. To mitigate these failures, we introduce a dense, decision-level reward in which an LLM judge evaluates the necessity of each tool call. This tool-necessity reward effectively suppresses cue-driven tool use while preserving task performance, providing a practical approach to improving the robustness of LLM agent tool-use policies.
Yiwei Yang, Haoxiang Zhang, Bingbing Wen +6
University of Washington · University of California San Diego · Stanford University
Reinforcement learning (RL) trained language model agents with tool access are increasingly deployed in coding assistants, research tools, and autonomous systems. We introduce the Reward Hacking Benchmark (RHB), a suite of multi-step tasks requiring sequential tool operations with naturalistic shortcut opportunities such as skipping verification steps, inferring answers from task-adjacent metadata, or tampering with evaluation-relevant functions. RHB supports independent and chained task regimes, where chain length acts as a proxy for longer-horizon agent behavior. We evaluate 13 frontier models from OpenAI, Anthropic, Google, and DeepSeek. Exploit rates range from 0% (Claude Sonnet 4.5) to 13.9% (DeepSeek-R1-Zero), varying sharply by post-training style. A controlled sibling comparison (DeepSeek-V3 vs. DeepSeek-R1-Zero) shows RL post-training is associated with substantially higher reward hacking (0.6% vs. 13.9%), with consistent gaps across all four task families. We identify six exploit categories and find that 72% of reward hacking episodes include explicit chain-of-thought rationale, suggesting models often frame exploits as legitimate problem-solving. Simple environmental hardening reduces exploit rates by 5.7 percentage points (87.7% relative) without degrading task success. Models with near-zero exploit rates on standard tasks show elevated rates on harder variants, suggesting that production-aligned post-training appears to suppress reward hacking only below a complexity threshold where honest solutions remain tractable.
Agentic reinforcement learning can induce tool abuse, where models overuse external tools even for queries solvable by internal reasoning. Existing approaches mitigate this issue with uniform tool-use penalties or hard limits, which reduce tool frequency but may also suppress useful tool-assisted exploration. We propose EAPO, an Efficient Agentic Policy Optimization framework that learns selective tool use. EAPO introduces tool-free trajectories into each rollout group, applies difficulty-aware reward shaping to penalize redundant tool calls mainly on easier queries, and uses confidence-aware token reweighting to improve policy learning. Across nine mathematical and knowledge-intensive reasoning benchmarks, EAPO consistently improves the accuracy efficiency trade-off on Qwen2.5-3B, Qwen2.5-7B, and Llama3.1-8B. Compared with GRPO, EAPO improves average performance by 10.45%, 7.27%, and 9.69%, while reducing average tool calls by 18.33%, 18.33%, and 24.59%, respectively. These results show that agents can learn when not to use tools without compromising tool-integrated reasoning.
Liuji Chen, Dianxing Tang, Xing Shi +4
1NLPR, Institute of Automation, Chinese Academy of Sciences · 3Zhejiang University · 2ByteDance