Organizations: Mohamed bin Zayed University of Artificial Intelligence · University of Electronic Science and Technology of China · Nanjing University · Griffith University · University of New South Wales · King’s College London
Tool calling, invoking external tools on demand, is central to agentic LLMs, yet the mechanism that decides whether a model calls a tool or responds directly remains poorly understood. Agentic prompts are long and heavily scaffolded, combining role instructions, tool schemas, format templates, and the user's request across hundreds of tokens, creating a noisy, highly entangled context in which no single controllable variable for mechanistic analysis is obvious. To obtain such a variable, we propose a method that converts complex agentic prompts into minimal contrastive pairs in which a single request verb determines the tool-call decision: replacing an execution-verb (e.g., \textit{write}) with an analysis-verb (e.g., \textit{discuss}) reliably flips the decision, suggesting it is mediated by a compact internal state. We construct 500 such paired prompts across Python, Java, and C++ (300 for mechanistic analysis, 200 held out for evaluation). We trace the decision to a vector, μΔ, that is both causally necessary and sufficient and generalizes beyond the discovery prompts to native multi-turn τ2-Bench trajectories and verb-free requests. Behavioral ablations show that the scaffold establishes a tool-call prior; Transcoder decomposition then reveals that analysis verbs suppress this prior through features signaling that tool use is unnecessary, whereas execution verbs largely leave it intact. Downstream scaffold-reading attention heads and MLP features read out the resulting state, and the same mechanism recurs across seven models from the Qwen, Mistral, and Granite families. Our code is available at https://github.com/XijieGo/MI4ToolCalling.
Figures & tables
Figure 1: Roadmap of the paper’s mechanistic analysis. The three numbered blocks summarize the stages of our analysis: (1) verb substitution turns long agentic prompts into a controlled causal interface; (2) the call-or-no-call decision passes through a tool-call vector, μΔ , that is causally sufficient and necessary; and (3) this vector reflects suppression of a scaffold-induced tool-call default and generalizes across the model families. The six lettered steps show how the evidence is built.
Figure 2: Request verbs change tool-call rates across models. Bars show first-token call rates on HumanEval before behavioral filtering, by request type.
Figure 3: The call signal shifts from the verb to the prediction position. Layerwise patching finds the L24 residual state that determines when to call a tool or to give a text reply.
Intervention
Top-1 before
Top-1 after
Δ top-1
Logit before
Logit after
Δ logit
Suff.
Necc.
Add μΔ
0.00%
100.00%
{\color[rgb]{1,0,0}\uparrow}\,100.00
25.11
32.50
{\color[rgb]{1,0,0}\uparrow}\,7.39
1.03
–
Remove μΔ
100.00%
0.00%
{\color[rgb]{0,0.6,0}\downarrow}\,100.00
32.28
24.80
{\color[rgb]{0,0.6,0}\downarrow}\,7.48
–
1.04
Table 1: The tool-call vector causally controls the call-or-no-call decision. Adding μΔ to analysis prompts induces tool calls, while removing it from execution prompts suppresses them.
Add μΔ
Remove μΔ
Target
Tool
N
Induction (%)
Norm. shift
Suppression (%)
Norm. shift
Web retrieval
web_search
100
100.0
0.93
100.0
0.75
SQL execution
run_sql
100
100.0
0.79
100.0
0.77
Email dispatch
send_email
100
100.0
0.91
100.0
1.30
Mean
—
300
100.0
0.88
100.0
0.94
Table 2: The coding-derived vector controls tool calling across web retrieval, SQL execution, and email dispatch. Qwen3-8B interventions keep the coding-derived direction fixed at L24 and match the vector norm to each target domain. N counts held-out pairs. Induction and suppression rates measure first-token switches from no-call to call under addition and from call to no-call under removal, respectively. Norm. shift is the mean <tool_call> logit increase for Add or decrease for Remove, divided by the baseline call–no-call logit gap. Means weight domains equally.
Setting
Intervention
N
Top-1 flip (%)
Mean Δzcall
Top-1 rate (before → after)
τ2 -Bench, native call
Remove
200
36.7
−5.55
99.5% → 63.0%
τ2 -Bench, native text
Add
200
70.0
+19.42
0.0% → 70.0%
Verb-free, native call
Remove
160
93.1
−9.95
100.0% → 6.9%
Table 3: The coding-derived Qwen3-8B vector transfers beyond the discovery prompts. All entries use 1× gain at L24. N counts evaluated decision points. Top-1 flip is the fraction of baseline calls suppressed by removal or baseline text responses changed to calls by addition. Δzcall is the mean paired after-minus-before change in <tool_call> logit over all N points. Top-1 rate gives the call rate before and after intervention over the same N points. Appendix C gives random controls and shows how the verb-free results vary with the chosen intervention gain.
Scaffold
Neutral request
Analysis request
R+T+F (full)
0.8486
0.0038
R+T (no format template)
6.32×10−7
8.31×10−10
R+F (no tool schema)
0.9521
0.4108
F only
1.0000
0.9952
Table 4: The scaffold creates a request-sensitive call prior. Entries are mean first-token call probabilities on 200 held-out Qwen3-8B tasks. Appendix D.1 gives all component combinations and a control that replaces the tool schema with text of the same length.
Figure 4: MLPs dominate formation-window writes. Left, Qwen3-8B component outputs projected onto μ^Δ . Right, the Transcoder architecture used to decompose MLP writes.
Layer
Dominant
Kcorrupt
Kclean
Kcorrupt/Kclean
Share (%)
Semantic label
L20
Clean
0.24
0.93
0.26
4.6
Execution requests
L21
Corrupt
1.47
1.10
1.34
16.7
Non-necessity
L22
Corrupt
1.85
1.33
1.39
22.8
Analysis-task contexts
L23
Corrupt
5.98
1.43
4.18
55.9
Analysis-verbs
Table 5: Features more active on analysis prompts dominate the formation window. Kcorrupt and Kclean sum ∣κ∣ over all features with higher analysis and execution activations, respectively, on 200 held-out pairs. Share is each layer’s fraction of the net Transcoder feature write over L20–L23. Appendix D.4 gives the statistics for the labelled families of Transcoder features.
Figure 5: Scaffold-reading heads and a late structural feature read out the tool-call state. (A) Execution-minus-analysis DLA for the <tool_call> logit across Qwen3-8B attention heads in L25–L35. (B) L29H9, L33H11, and L33H29 share a shift toward the tool-format template and away from role instructions on execution prompts (blue) relative to analysis prompts (red). (C) Max-activating contexts for L34/F109925 highlight tool-schema boundaries, supporting its interpretation as a structural readout feature for the opening of a tool call.
Localization
Intervention
Generalization (%)
Formation
Readout
Model
l
r(l,p) (%)
Suff.
Necc.
Domains
τ2
Verb-free
MLP/Attn
Kcorrupt/Kclean
Max Attn (pp)
Qwen3-4B
26
100.0
0.97
0.71
87.7
72.2
100.0
2.08
3.26
82.6
Qwen3-8B
24
100.0
1.03
1.04
100.0
53.3
93.1
3.48
1.99
72.6
Qwen3-14B
34
100.0
0.97
0.84
100.0
85.0
100.0
6.87
2.41
87.3
Qwen3.5-4B
31
99.0
0.84
0.61
91.1
47.9
100.0
–
1.58
50.3
Qwen3.5-9B
31
100.0
0.92
0.88
100.0
80.2
100.0
–
1.30
52.2
Table 6: The tool-call mechanism recurs across seven models and generalizes beyond the discovery prompts. Domains and τ2 average conditional intervention rates; Verb-free reports suppression at unit gain. Domains uses norm calibration (Appendix I.1 ). Max Attn summarizes clean-minus-corrupt scaffold-attention shifts (Appendix F.3 ). For hybrid Qwen3.5, MLP/Attn is omitted and attention is summarized only for layers with full attention.
Programming tasks with statements, function signatures, examples, tests, or reference solutions.
τ2 -Bench Telecom
Qwen3.5-4B, Mistral-Small-3.2-24B
Conversation histories with a controlled request-verb contrast.
Appendix
Table 7: Source families represented in the model-specific paired corpora.
Model
Source family
Train
Held-out
Qwen3-4B
Coding tasks
300
200
Qwen3-8B
Coding tasks
300
200
Qwen3-14B
Coding tasks
300
200
Qwen3.5-4B
τ2 -Bench Telecom
300
200
Qwen3.5-9B
Coding tasks
300
200
Mistral-Small-3.2-24B
τ2 -Bench Telecom
300
200
Appendix
Table 8: Per-model paired-prompt corpus sizes and splits.
Split
Execution verbs
Analysis verbs
Train (300)
add 60, build 60, complete 60, save 60, write 60
discuss 75, explore 75, review 75, study 75
Held-out (200)
add 53, build 20, complete 20, save 53, write 54
discuss 40, explore 40, inspect 40, review 40, study 40
Appendix
Table 9: Qwen3-8B execution- and analysis-verb counts by split.
Source
Train
Held-out
APPS
188
51
CodeContests
63
135
HumanEval
17
4
MBPP
32
10
Total
300
200
Appendix
Table 10: Source composition of the Qwen3-8B paired-prompt corpus.
Intervention
L22
L23
L24
L25
Analysis +μΔ
82.0%
100.0%
100.0%
100.0%
Execution −μΔ
45.7%
12.0%
0.0%
0.0%
Analysis + random
0.0%
0.3%
18.7%
0.0%
Execution − random
100.0%
99.7%
97.7%
100.0%
Appendix
Table 11: Neighboring-layer control with independently fitted directions. Entries are held-out post-intervention call rates. Random controls match the learned direction’s norm at each layer.
Model
μΔ , 1×
Random, 1×
Qwen3-4B
61.3%
0.5%
Qwen3-8B
36.7%
1.5%
Qwen3-14B
97.0%
8.0%
Qwen3.5-4B
88.4%
0.0%
Qwen3.5-9B
91.3%
17.3%
Mistral-Small-3.2-24B
75.0%
0.5%
Appendix
Table 12: τ2 -Bench removal: strict first-token flip rate among native tool-call decisions ( Ncall baseline tool calls per model; N=199 for Qwen3-8B).
Model
μΔ , 1×
Random, 1×
Qwen3-4B
83.0%
0.0%
Qwen3-8B
70.0%
0.0%
Qwen3-14B
73.0%
0.0%
Qwen3.5-4B
7.5%
0.0%
Qwen3.5-9B
69.0%
0.0%
Mistral-Small-3.2-24B
92.0%
0.0%
Appendix
Table 13: τ2 -Bench induction: strict first-token flip rate among native text decisions.
Model
N
(−μΔ) , α=1
(−μΔ) , α=1.5
(−r) , α=1.5
Qwen3-4B
117
100.0%
100.0%
34.2%
Qwen3-8B
160
93.1%
100.0%
16.2%
Qwen3-14B
161
100.0%
100.0%
44.7%
Qwen3.5-4B
186
100.0%
100.0%
33.9%
Qwen3.5-9B
187
100.0%
100.0%
73.8%
Mistral-Small-3.2-24B
147
100.0%
100.0%
5.4%
Appendix
Table 14: Implicit, verb-free generalization across all seven models. N is the evaluated baseline-positive subset; r is a norm-matched random direction.
Scaffold components
Neutral pcall
Analysis pcall
P
R+T+F (full scaffold)
0.8486
0.0038
0.845
R+F (no tool schema)
0.9521
0.4108
0.541
R+T (no format template)
6.32×10−7
8.31×10−10
6.3×10−7
F only
1.0000
0.9952
0.005
T only
1.87×10−8
2.05×10−11
1.9×10−8
R only
8.77×10−9
7.14×10−11
8.7×10−9
Appendix
Table 15: Scaffold-component ablation, Qwen3-8B, 200 held-out prompts per condition. pcall is the mean first-token <tool_call> probability; P is the neutral-minus-analysis gap.
Request
Call top-1
Mean call probability
Empty turn, greeting, or unrelated question
0.0%
≈5.0×10−4 (max.)
Task body only
32.3%
0.3310
Neutral verb and task body
86.7%
0.8547
Analysis verb and task body
0.0%
0.0080
Execution verb and task body
100.0%
0.9993
Execution verb and task body, no scaffold
0.0%
6.56×10−17
Appendix
Table 16: Qwen3-8B request baselines on a 300-task evaluation cohort. Empty, greeting, and unrelated turns are evaluated once each. The full scaffold is present except in the final row.
Request verb
write_file
submit_review
Write
100.0%
34.0%
Review
4.0%
100.0%
Appendix
Table 17: Request–tool crossover on Qwen3-8B. Entries are first-token call rates, with 100 requests per condition.
Diagnostic
Result
Probe AUC >0.9 first layer
L1
Logit-lens gap >1 first layer
L27
DLA top-3 layers (MLP write)
L35, L34, L33
Final gap mean (single RMSNorm, 95% CI)
7.093 [6.687, 7.498]
First 10% / 25% / 50% of final gap
L25 / L29 / L29
Largest increments
L29, L33, L24, L27, L32
Appendix
Table 18: Held-out early-signal diagnostics and layer-wise gap accumulation on the 200 held-out pairs (single-sample TransformerLens protocol with standard single RMSNorm).
Family
Side
Nf
Mean
Sum
Interpretation
analysis_non_execution
corrupt
115
0.0633
7.2787
Analysis requests and non-necessity semantics.
execution_request
clean
292
0.0156
4.5634
Execution requests and action verbs.
analysis_plain_discourse
corrupt
15
0.0020
0.0294
General analysis discourse.
schema_boundary_formatting
corrupt
4
0.0466
0.1862
Code and format boundaries.
Appendix
Table 19: Formation-window feature families selected on training pairs. Side is the majority activation side on training pairs. Counts cover 426 of 640 candidates; 214 remain unlabelled. Mean and sum use held-out ∣κ∣ .
Layer
Kcorrupt
Kclean
Ratio
Net-write share
L20
0.24
0.93
0.26
4.6%
L21
1.47
1.10
1.34
16.7%
L22
1.85
1.33
1.39
22.8%
L23
5.98
1.43
4.18
55.9%
Appendix
Table 20: Layer-wise formation on the 200 Qwen3-8B feature-evaluation held-out pairs. Kcorrupt and Kclean sum ∣κ∣ over features with higher activation on the corresponding side. Net-write share is the layer’s sum of κ divided by the total over L20–L23.
Execution prompts
Analysis prompts
Condition
gl
Call (%)
gl
Call (%)
Baseline
60.06
100.0
60.06
0.0
MLP-output swap
15.78
93.0
15.09
95.0
Swap + μ^Δ restoration
60.06
100.0
60.06
0.0
Swap + orthogonal restoration
15.78
58.0
15.09
99.5
Swap + full-state restoration
60.06
100.0
60.06
0.0
Appendix
Table 21: Formation-window mediation on the 200 feature-evaluation pairs. Each arm computes gl against the untouched paired opposite-side state. Call rate is the post-intervention top-1 rate.
Head
stage
DLA Δ
recover shift
corrupt top-1
strict flip
L25H6
earlier
0.06
1.28
40.4%
36.9%
L28H3
earlier
0.26
1.38
36.1%
32.5%
L26H14
earlier
0.04
0.91
24.3%
20.7%
L29H9
earlier
4.15
1.34
21.3%
17.8%
L33H29
late
7.91
0.62
8.5%
4.9%
L33H11
late
3.32
-0.39
2.4%
0.0%
Appendix
Table 22: Selected downstream attention heads in the L25–L35 component sweep ( N=300 candidate pairs; baseline corrupt tool-call rate is 3.6%). DLA delta is clean minus corrupt for the <tool_call> logit. Figure 5A’s heatmap displays this component sweep cohort (with L33H29 DLA Δ=7.91 ). In contrast, Section 6.1 and Table 24 report on the dedicated 200 balanced held-out pairs (where L33H29 exhibits Clean DLA 24.58 and Corrupt DLA 11.13, giving Δ=13.45 ). Recovery shift and strict flip are measured by replacing the corrupt-side head output with the clean-side output.
MLP
recover shift
drop shift
corrupt top-1
strict flip
L34
1.74
2.32
53.3%
49.7%
L27
1.32
0.82
52.3%
48.7%
L25
1.70
0.96
44.8%
41.2%
L35
0.84
1.27
31.8%
28.2%
L31
0.90
0.90
28.6%
25.0%
L29
0.38
1.58
26.6%
23.1%
Appendix
Table 23: Top downstream MLP blocks in the L25–L35 component sweep ( N=300 , baseline corrupt top-1 call rate 3.6%), separate from the 200-pair prediction-position patching subset (42.5% top-1 / 39.0% strict recovery) and the 200-pair feature-evaluation cohort (41.0% top-1). Each row patches only one MLP output from clean into corrupt, or from corrupt into clean for the reverse drop measurement.
Head
Clean
Corrupt
Corrupt +μΔ
L29H9
6.99
1.43
6.91
L33H11
5.60
1.46
5.30
L33H29
24.58
11.13
24.43
Appendix
Table 24: Attention-head response to L24 vector addition on the 200 balanced Qwen3-8B held-out pairs. Entries are mean DLA for the call-opening token.
Measurement
Clean
Corrupt
Corrupt +μΔ
Mean activation
124.43
67.63
126.00
Projected call-token write
6.17
3.35
6.25
Appendix
Table 25: L34 F109925 on the 200 Qwen3-8B feature-evaluation held-out pairs. Projected write is activation multiplied by the feature decoder’s projection onto the call-token unembedding direction.
Condition
Call (%)
Corrupt baseline
0.0
Corrupt +μΔ at L24
100.0
Replace F109925
37.0
Replace F91365
6.5
Patch full MLP34
41.0
Appendix
Table 26: L34 feature replacement and controls on the same 200 held-out pairs. All interventions act at the prediction position.
Model
Ltot
L∗
r(l,p) (%)
Suff.
Necc.
Qwen3-4B
36
L26
100.0
0.97
0.71
Qwen3-8B
36
L24
100.0
1.03
1.04
Qwen3-14B
40
L34
100.0
0.97
0.84
Qwen3.5-4B
32
L31
99.0
0.84
0.61
Qwen3.5-9B
32
L31
100.0
0.92
0.88
Mistral-Small-3.2-24B
40
L25
100.0
0.90
0.76
Appendix
Table 27: Per-model localization and shared-direction results. L∗ is the commitment layer; r(l,p) is the top-1 native call-token recovery rate at the prediction position. Suff. and Necc. are the causal sufficiency and necessity scores for the shared direction μΔ .
Model
Ltot
L∗
N
Probe AUC >0.9
Gap >1.0
Top-3 DLA layers
Qwen3-4B
36
L26
300
L0
L26
L34, L33, L32
Qwen3-8B
36
L24
200
L1
L27
L34, L33, L29
Qwen3-14B
40
L34
300
L1
L25
L34, L37, L38
Qwen3.5-4B
32
L31
300
L0
L14
L31, L30, L27
Qwen3.5-9B
32
L31
300
L0
L21
L27, L30, L23
Mistral-Small-3.2-24B
40
L25
300
L0
L17
L39, L37, L34
Appendix
Table 28: Early-signal and timing diagnostics across all seven models on held-out prompt pairs. L∗ is the causal commitment layer from Table 6 ; Probe AUC >0.9 is the first layer where the linear probe achieves >0.9 AUC; Gap >1.0 is the first layer where the logit-lens clean-minus-corrupt gap exceeds 1.0; Top-3 DLA reports the layers with the largest signed positive net direct logit attribution differences ( ΔDLA>0 ). L∗ denotes a block input; probe and lens layer labels denote block-output residuals, and DLA labels denote block writes.
Figure 6: Triplet diagnostics across all seven models, arranged with one model per row and three diagnostics per row. From left to right, each row shows the linear probe AUC, the logit-lens clean-minus-corrupt gap, and the layer-wise DLA sweep. Evaluation cohorts contain 200 pairs for Qwen3-8B, 240 for Granite-3.3-8B, and 300 for each of the other five models. Across all seven models, probe separability appears at L0–L1, while the largest DLA writes occur in upper layers.
Model
MLP/Attn
Kcorrupt/Kclean
Qwen3-4B
2.08
3.26
Qwen3-8B
3.48
1.99
Qwen3-14B
6.87
2.41
Qwen3.5-4B
–
1.58
Qwen3.5-9B
–
1.30
Mistral-Small-3.2-24B
3.93
1.61
Appendix
Table 29: Formation-window summary across models. Kcorrupt/Kclean is the ratio of corrupt-higher to clean-higher feature mass. Qwen3.5-4B feature mass is evaluated over L28–L30.
Model
Max Attn (pp)
Qwen3-4B
82.6
Qwen3-8B
72.6
Qwen3-14B
87.3
Qwen3.5-4B
50.3
Qwen3.5-9B
52.2
Mistral-Small-3.2-24B
40.5
Appendix
Table 30: Cross-model scaffold-attention summary on native held-out pairs. Max Attn reproduces Table 6 , in percentage points. Attention shifts are averaged over held-out pairs, then maximized over measured heads. The shift is clean minus corrupt in the target scaffold region. Qwen3-8B uses its three reported scaffold readers; Qwen3.5 uses full-attention layers only.
Model
Core nodes
Core edges
Suff.
Removal
Rand. Suff.
Gap
Qwen3-1.7B
18
54
0.9891
0.9824
0.8812
0.1078
Qwen3-4B
18
69
0.9979
0.8319
0.8712
0.1266
Qwen3-8B
18
63
1.0000
0.5975
0.9768
0.0232
Qwen3-14B
20
86
0.9999
0.4077
0.9554
0.0446
Appendix
Table 31: Sparse forward-core replay across Qwen3 scales. Core nodes and edges are taken from the aggregated forward sparse graph. Suff. is the mean replay ratio of the discovered core relative to the clean baseline, Removal is the mean ratio under clean-run removal, Rand. Suff. is the mean replay ratio of size-matched random cores, and Gap is the mean advantage of the discovered core over the random baseline.
Figure 7: EAP-IG top- k faithfulness curve for the tool-call task, normalized between the corrupt and clean baselines. The curve relates retained edge budget to recovery of the <tool_call> signal.
Retained edges
Faith.
Suff. logit diff
Nodes
Core overlap
1,000
-0.005
-13.12
139
17/18
2,000
-0.014
-13.38
229
18/18
5,000
0.043
-11.81
417
18/18
10,000
0.284
-5.19
553
18/18
20,000
0.451
-0.59
708
18/18
30,000
0.470
-0.08
805
18/18
Appendix
Table 32: EAP-IG top- k sweep on Qwen3-8B using 100 equal-length clean/corrupt prompt pairs. Faith. normalizes sufficiency between the corrupt and clean baselines. Suff. is the mean logit-difference score under corrupt-side replay with only the retained edges active. Nodes counts the retained component set induced by the edges; core overlap counts shared nodes with the 18-node forward core.
Figure 8: Prompt-specific feature graph from Circuit Tracing [ 28 ] . Nodes represent features or residual-stream components, and edges represent direct-effect attribution weights. The graph describes local contributors to the tool-call computation.
Source benchmark
License
APPS
MIT
HumanEval
MIT
MBPP
CC BY 4.0
CodeContests
CC BY 4.0 (data); Apache 2.0 (code)
τ2 -Bench
MIT
FEVER
CC BY-SA 3.0 (data)
Appendix
Table 33: Licenses of the source benchmarks. Links identify the original resources; the original papers are cited in the dataset and transfer sections.
Tool-augmented LLM agents tend to call tools indiscriminately, even when the model can answer directly. Each unnecessary call wastes API fees and latency, yet no existing benchmark systematically studies when a tool call is actually needed. We propose When2Tool, a benchmark of 18 environments (15 single-hop, 3 multi-hop) spanning three categories of tool necessity -- computational scale, knowledge boundaries, and execution reliability -- each with controlled difficulty levels that create a clear decision boundary between tool-necessary and tool-unnecessary tasks. We evaluate two families of training-free baselines: Prompt-only (varying the prompt to discourage unnecessary calls) and Reason-then-Act (requiring the model to reason about tool necessity before acting). Both provide limited control: Prompt-only suppresses necessary calls alongside unnecessary ones, and Reason-then-Act still incurs a disproportionate accuracy cost on hard tasks. To understand why these baselines fail, we probe the models' hidden states and find that tool necessity is linearly decodable from the pre-generation representation with AUROC 0.89--0.96 across six models, substantially exceeding the model's own verbalized reasoning. This reveals that models already know when tools are needed, but fail to act on this knowledge during generation. Building on this finding, we propose Probe&Prefill, which uses a lightweight linear probe to read the hidden-state signal and prefills the model's response with a steering sentence. Across all models tested, Probe&Prefill reduces tool calls by 48% with only 1.7% accuracy loss, while the best baseline at comparable accuracy only reduces 6% of tool calls, or achieves a similar tool call reduction but incurs a 5× higher accuracy loss. Our code is available at https://github.com/Trustworthy-ML-Lab/when2tool
Agentic AI architectures augment LLMs with external tools, unlocking strong capabilities. However, tool use is not always beneficial; some calls may be redundant or even harmful. Effective tool use, therefore, hinges on a core LLM decision: whether to call or not call a tool when performing a task. This decision is particularly challenging for web search tools, where the benefits of external information depend on the model's internal knowledge and its ability to integrate potentially noisy tool responses. We introduce a principled framework inspired by decision-making theory to evaluate web search tool-use decisions along three key factors: necessity, utility, and affordability. Our analysis combines two complementary lenses: a normative perspective that infers true need and utility from an optimal allocation of tool calls, and a descriptive perspective that infers the model's self-perceived need and utility from their observed behaviors. We evaluate six open and one closed-source frontier models under two harnesses, one conditioning on only the current turn and its search results, the other on the full execution traces, across four web-search tools and three tasks. In every setting, we find that a model's perceived need and utility are frequently misaligned with the true need and utility. Building on this framework, we train lightweight estimators of need and utility from the models' hidden states. These estimators drive simple controllers that improve decision quality and yield stronger task performance than the self-perceived baseline for most of the open-source models.
Qinyuan Wu, Soumi Das, Mahsa Amani +5
Max Planck Institute for Software Systems · Ruhr University Bochum · UAR RC Trust
When a tool-calling agent picks the wrong tool, the failure is invisible until execution: the email gets sent, the meeting gets missed. As agents take on consequential actions, one bad tool call can do real damage. We currently have no way to look inside the model and catch the mistake before it happens; this paper shows that we can. Inside the model, the choice of tool is carried by a single direction in activation space, one direction per pair of tools. Adding that direction during generation switches which tool the model picks. Across 12 instruction-tuned and 6 base models spanning Gemma 3, Qwen 3, Qwen 2.5, and Llama 3.1 (270M to 27B), this works at 83-100% accuracy on 4B+ instruction-tuned models on a 15-tool synthetic benchmark and at 77-94% on the real-API benchmark τ-bench airline. The JSON arguments that follow automatically adapt to the new tool's schema, so flipping the name is enough. The same per-tool directions also flag likely errors before they happen: queries where the model is unsure between two tools fail 21x more often than queries where it is not (Gemma 3 27B). This is not just topic injection: random vectors at the same magnitude give a 0% switch rate, and a probe within a single domain (14 airline tools that share one topic) still reads which tool the model will call at top-1 61-89% across five 4B-14B models. Even base models already carry the right tool internally before they can emit it: reading the chosen tool off the model's internal state (cosine readout) recovers 61-82% accuracy on BFCL while base generation lands at 2-10%, suggesting pretraining forms the representation and instruction tuning later wires it to the output. Our results cover single-turn, fixed-menu settings; on multi-turn agent loops the same intervention is less stable (matched-baseline gain or loss of up to 30 percentage points with no consistent direction).
Zekun Wu, Ze Wang, Seonglae Cho +4
University College London · Holistic AI · Imperial College London