Reasoning language models that can call tools must decide during inference whether to answer unaided or delegate. Any self-reflection mechanism for this must answer three questions: where the reflective signal comes from (verbal reports, output distributions, hidden states, a separate predictor), how it is presented to the model (numerical prediction, confidence token, prompt injection), and whether it changes the model's subsequent action. We isolate the third question. At a fixed point in otherwise identical reasoning trajectories, we insert a single first-person sentence expressing either confidence or doubt; the model then continues reasoning and chooses whether to answer directly or call a tool. Comparing these counterfactual continuations measures the causal effect of the reflective signal on delegation. We call this behavioral response Nudgeability and measure it along two dimensions: sensitivity, how strongly confidence and doubt change delegation rates, and targeting, whether delegation increases for problems the model cannot solve unaided and decreases for those it can. Across nine small-to-medium open-weight reasoning models from three families (Qwen, Gemma, and GLM) and two tasks, models are consistently sensitive: doubt increases delegation and confidence decreases it, with a median confidence-to-doubt swing of 20.6 percentage points, and 53 to 70 points for the larger provider-served models. This responsiveness is poorly targeted: a median 42% of induced flips are well-targeted, only a +2 percentage-point lift over a random-selection baseline. Confidence language is thus a strong control surface for delegation, but current models use it only weakly in accordance with their actual competence. Nudgeability offers a simple, post-training-free way to evaluate both sensitivity and targeting as endogenous self-reflection mechanisms mature.
Figures & tables
MuSiQue 2-hop
StrategyQA
Model
S (pp)
WT (%)
L (pp)
S (pp)
WT (%)
L (pp)
Open-weight models ( n=1000 )
Qwen3-4B
14.0
58.2
18.8 †
22.2
42.3
1.4
Qwen3-8B
22.4
42.5
1.9
19.0
39.5
3.8 †
Qwen3-14B
31.0
41.7
5.2 †
33.9
49.2
2.9
Qwen3-32B
38.3
37.2
6.8 †
38.3
54.7
1.7
Table 1: Confidence–doubt delegation swing ( S ), well-targeted flip share ( WT ), and targeting lift ( L ) across both datasets. Nominal n=1000 for open-weight models and n=250 for provider-served models. Point estimates are reported here; † marks a lift whose 95% problem-bootstrap interval excludes zero, and the median row is taken per task over the nine open-weight models. Full 95% pointwise problem-bootstrap intervals, retained sample sizes ( nS/nWT ), flip event counts (H/F), and random baselines appear in Appendix G.1 .
Appendix figures & tables8 assets
Supplementary material from the paper’s appendix.
Appendix
MuSiQue 2-hop
StrategyQA
Model
correct/known
Acc. (%)
correct/known
Acc. (%)
Open-weight models
Qwen3-4B
521/889
58.6
711/991
71.7
Qwen3-8B
597/939
63.6
763/999
76.4
Qwen3-14B
625/990
63.1
757/999
75.8
Qwen3-32B
658/992
66.3
816/1000
81.6
Appendix
Table 2: Unaided accuracy of the no-tool base, P(s=1) : correct answers over non-truncated no-tool generations (truncated generations have unknown competence and are excluded). Nominal n=1000 for open-weight models and n=250 for provider-served models.
Task
Model
nS/nWT
S [CI]
H/F
WT [CI]
WTrand [CI]
Lift [CI]
Open-weight models ( n=1000 )
MuSiQue
Qwen3-4B
966/876
14.0 [11.8, 16.3]
78/134
58.2 [49.2, 67.2]
39.4 [35.7, 43.1]
18.8 [10.7, 27.1]
MuSiQue
Qwen3-8B
944/921
22.4 [19.6, 25.1]
128/301
42.5 [36.4, 48.8]
40.6 [37.1, 44.1]
1.9 [-3.1, 7.1]
MuSiQue
Qwen3-14B
990/986
31.0 [28.0, 34.0]
134/321
41.7 [36.4, 47.1]
36.5 [33.5, 39.6]
5.2 [0.9, 9.6]
MuSiQue
Qwen3-32B
975/981
38.3 [35.1, 41.3]
136/366
37.2 [32.1, 42.2]
30.4 [27.4, 33.3]
6.8 [3.0, 10.7]
MuSiQue
Gemma4-E2B
1000/993
24.2 [21.5, 27.0]
108/293
36.9 [31.1, 42.9]
44.7 [41.1, 48.3]
-7.9 [-12.6, -3.0]
Appendix
Table 3: Confidence–doubt delegation swing ( S ), well-targeted flip share ( WT ), random-selection reference ( WTrand ), and targeting lift for every experiment, open continuation. S and lift are in percentage points; WT and WTrand are percentages. H/F gives helpful/all flip events. Brackets are pointwise 95% problem-bootstrap intervals. Appendix G.1 describes the columns and retained populations.
Model
Task
S0
S1
S2
S3
Spool [95% CI]
Qwen3-8B
StrategyQA
24.2
19.4
14.0
21.6
18.2 [13.1, 23.7]
Qwen3-8B
MuSiQue
21.4
17.5
11.2
15.5
14.7 [9.1, 20.6]
Gemma4-E4B
StrategyQA
24.2
18.6
21.3
22.7
20.8 [15.5, 26.3]
Gemma4-E4B
MuSiQue
70.0
40.4
45.1
43.5
43.0 [36.2, 49.8]
GLM-Z1-9B
StrategyQA
8.1
14.1
11.1
15.2
13.5 [9.1, 18.5]
GLM-Z1-9B
MuSiQue
3.4
5.0
6.0
1.0
4.0 [1.7, 6.7]
Appendix
Table 4: Stochastic-decoding sensitivity at temperature 1.0 and top- p=0.95 . S0 is the greedy estimate; S1 – S3 are decode-seed estimates. Spool pools seed-paired problem contributions, and brackets give a 95% problem-clustered bootstrap interval. All values are percentage points.
Unaided accuracy
WT
Model
Task
0
1
2
3
0
1
2
3
pool [95% CI]
Qwen3-8B
StrategyQA
72.0
74.0
75.0
75.0
52.2
32.1
43.5
29.6
34.6 [22.8, 47.9]
Qwen3-8B
MuSiQue
64.9
64.0
66.0
62.0
51.5
58.5
56.2
27.0
47.3 [35.9, 58.4]
Gemma4-E4B
StrategyQA
70.0
68.0
67.0
63.0
36.0
66.7
63.9
40.5
56.0 [44.9, 67.0]
Gemma4-E4B
MuSiQue
62.0
62.0
67.0
65.0
25.0
21.6
34.0
35.0
29.7 [20.4, 39.8]
GLM-Z1-9B
StrategyQA
72.7
71.4
68.0
70.7
50.0
68.8
26.3
60.0
50.0 [36.1, 64.1]
Appendix
Table 5: Unaided accuracy and well-targeted share under stochastic decoding (temperature 1.0, top- p=0.95 , nominal n=100 ). Subscript 0 is the greedy run and 1–3 are decode seeds; each seed relabels competence from its own sampled no-tool answers. WTpool pools seed-paired flip contributions, with a 95% problem-clustered bootstrap interval. All values are percent.
Task
Model
n
B
∅
ν0
ν+
ν−
MuSiQue
Qwen3-4B
955
10.7
9.8
9.3
8.5
22.3
Qwen3-8B
930
36.0
35.4
27.4
27.7
50.1
Qwen3-14B
990
7.5
6.3
6.6
4.7
35.8
Qwen3-32B
970
10.1
7.8
8.2
6.9
45.2
Gemma4-E2B
1000
10.0
10.0
5.6
4.5
28.7
Gemma4-E4B
1000
14.8
14.7
15.3
13.4
79.9
Appendix
Table 6: Delegation rate P(A=\textsctool) , in percent, for the uninterrupted with-tool base B and the four splice arms: sentence-less ( ∅ ), neutral ( ν0 ), confidence ( ν+ ), and doubt ( ν− ). Nominal n=1000 ; each row’s arms share one pairwise-complete denominator n .
Task
Model
Dself (count)
Rself (count)
MuSiQue
Qwen3-4B
10.0 (49/489)
12.2 (38/312)
Qwen3-8B
8.0 (33/411)
11.9 (21/177)
Qwen3-14B
9.5 (56/587)
12.3 (39/316)
Qwen3-32B
5.6 (34/607)
12.7 (32/251)
Gemma4-E2B
21.0 (99/471)
11.1 (47/423)
Gemma4-E4B
4.7 (25/527)
24.8 (54/218)
Appendix
Table 7: Supporting analysis at nominal n=1000 . Correctness transitions when the tool-aware base still answers for itself. Rates are percentages. Dself is correct-to-incorrect degradation and Rself is incorrect-to-correct recovery; parentheses give numerator/eligible retained-self denominator.
Task
Model
First person
Expert
Task
Model
First person
Expert
MuSiQue
Qwen3-4B
14.0
2.7
StrategyQA
Qwen3-4B
22.2
15.6
Qwen3-8B
22.4
12.8
Qwen3-8B
19.0
15.7
Qwen3-14B
31.0
11.4
Qwen3-14B
33.9
47.9
Qwen3-32B
38.3
13.8
Qwen3-32B
38.3
39.2
Gemma4-E2B
24.2
10.3
Gemma4-E2B
7.3
2.4
Gemma4-E4B
66.5
37.2
Gemma4-E4B
25.8
11.0
Appendix
Table 8: Supporting voice control at nominal n=1000 : open-continuation swing by voice endpoint, in percentage points.
Task
Model
Φ
ρ
ω
Task
Model
Φ
ρ
ω
MuSiQue
Qwen3-4B
-1.2
90.0
39.6
StrategyQA
Qwen3-4B
-17.9
44.1
24.5
Qwen3-8B
-1.8
64.3
29.0
Qwen3-8B
-2.8
22.7
18.9
Qwen3-14B
+0.3
92.2
34.1
Qwen3-14B
-20.2
43.5
17.5
Qwen3-32B
+2.0
87.9
29.4
Qwen3-32B
+0.2
18.9
11.1
Gemma4-E2B
+6.5
83.5
45.1
Gemma4-E2B
-12.9
33.8
34.2
Gemma4-E4B
+7.7
77.5
27.7
Gemma4-E4B
-7.5
41.8
20.5
Appendix
Table 9: Supporting oracle control at nominal n=1000 : decomposition in percent/percentage points. Negative Φ means the model delegates less to the generic free oracle than to the scoped tool.