Language models (LMs) often hallucinate by committing to confident answers rather than abstaining, even when they do not have enough information to answer reliably. A large body of existing work mitigates hallucination through detection or abstention mechanisms, but leaves open how models internally arrive at the decision to commit or abstain in the first place. We study this decision through mechanistic analysis, framing hallucination as unsupported commitment: the model commits despite exhibiting signals of unanswerability. Using causal gating, we identify a Commit-Abstain Circuit (CAC), a sparse, causally localised subset of attention heads and MLP sublayers underlying this decision. Across ten LMs (3B-14B) from five families and three benchmarks, the CAC exhibits a recurring accumulate-yet-undercorrect pattern: commitment-promoting components build up commitment in earlier layers, while abstention-promoting components act later as corrective signals that are often insufficient to overturn the accumulated commitment. Building on this finding, a lightweight policy trained on CAC activations improves decision accuracy by 12.2 points over the model's intrinsic commit-abstain margin, reduces false abstentions by 2.5 times, transfers to unseen benchmarks, and extends to larger models (27B-35B). The CAC is both diagnostic, clarifying how models overcommit, and practical, enabling improved abstention decisions.
Figures & tables
Figure 1: Unsupported commitment in Qwen 3.5 4B on an unanswerable KUQ question [ 27 ] . Left: Per-component contributions to the commit-abstain margin Δ(x) (Definition 1 ). Late-layer abstention signals fail to overturn earlier commitment ( Δ(x)=+3.25 ), leading to commitment. Right: A lightweight CAC-Informed MLP correctly triggers abstention.
Abstention set A
Commitment set C
“Don’t know”, “Insufficient evidence”, “Unsure”, “Unanswerable”, “Ambiguous”, “Beyond my knowledge”
All remaining, e.g. those initiating: “According to…”, “The answer is…”, “Exactly”, “Yes”, “No”, “In fact”
Table 1: Token sets used to define the commit-abstain margin Δ(x) .
Figure 3Figure 4Figure 5
Qwen 3.5
Gemma 3
Llama 3
Ministral
Phi-4
4B
9B
4B
12B
3B
8B
3B
14B
mini
14B
p
Mean AUROC
Zero-Threshold
.725
.798
.722
.854
.689
.784
.790
.848
.780
.830
6.0×10−6
Non-CAC Random
.691
.764
.741
.808
.648
.815
.732
.789
.756
.822
8.4×10−6
INSIDE
.572
.449
.586
.658
.578
.483
.406
.493
.415
.492
3.7×10−9
Multi-LLM
.688
.757
.683
.728
.538
.705
.596
.654
.603
.703
1.9×10−9
Semantic Entropy
.567
.463
.548
.600
.564
.521
.423
.564
.509
.529
1.3×10−8
Table 2: Mean AUROC and accuracy across three datasets for each model. Bold indicates the best result; underline indicates the second best. p -values are from two-sided Wilcoxon signed-rank tests over 30 model-dataset pairs. All differences are significant ( p<0.05 ), with most at p<10−5 .
Qwen 3.5
Gemma 3
Llama 3
Ministral
Phi-4
4B
9B
35B †
4B
12B
27B †
3B
8B
3B
14B
mini
14B
HotpotQA
Zero-Threshold
.887
.837
.867
.898
.922
.924
.840
.820
.851
.892
.859
.842
HaMI
.917
.938
.925
.894
.915
.892
.829
.924
.865
.889
.856
.937
CAC-Informed MLP (Ours)
.930
.949
.928
.911
.948
.950
.859
.913
.895
.894
.875
.939
SelfAware
Table 3: Decision accuracy on two unseen datasets (HotpotQA, SelfAware) and two unseen larger models ( † ). Our method achieves higher accuracy in all 24 configurations ( p=1.8×10−5 ).
Appendix figures & tables18 assets
Supplementary material from the paper’s appendix.
Appendix
Dataset
Source
Reasoning
Unanswerability
License
KUQ [ 27 ]
Parametric
Open-domain QA
Answers genuinely do not exist (unsolved problems, future events); tests limits of human knowledge.
MIT
SQuAD 2.0 [ 78 ]
In-context
Reading comprehension
Questions plausible but unsupported by the passage; tests resistance to surface-level cues.
CC BY-SA 4.0
MuSiQue [ 79 ]
In-context
Multi-hop
One supporting document removed ( missing evidence ); tests detection of a silent gap in the reasoning chain.
CC BY 4.0
SelfAware † [ 92 ]
Parametric
Self-knowledge
No definitive answer exists; tests limits of the model’s own knowledge, not of the world.
CC BY-SA 4.0
HotpotQA † [ 91 ]
In-context
Multi-hop
Supporting passages replaced by distractors ( misleading evidence ); tests resistance to plausible-but-wrong context.
CC BY-SA 4.0
Appendix
Table 4: Dataset summary. Source : whether the model must rely on parametric knowledge or a provided context passage. ( † ) held-out OOD evaluation sets only.
Figure 9: Annotation interface shown to Amazon Mechanical Turk workers for classifying candidate tokens as epistemic deferral or substantive answer.
I
“I don’t know”, “I cannot determine”
Sorry
“Sorry, I don’t have information…”
Unfortunately
“Unfortunately, I don’t have enough…”
Beyond
“Beyond my knowledge…”
Insufficient
“Insufficient information to answer…”
Impossible
“Impossible to determine without context”
Unanswerable
“Unanswerable based on context”
Without
“Without more context, I cannot…”
Unclear
“Unclear from the given context…”
Ambiguous
“Ambiguous; the premise is unclear”
Unknown
“Unknown; context does not specify…”
Unsure
“Unsure; evidence is inconclusive”
Uncertain
“Uncertain; lacks a definitive answer”
Unknowable
“Unknowable from the text provided”
Appendix
Table 5: Abstention token set A ( 17 tokens). The commitment set C contains all tokens not in A .
Qwen 3.5
Gemma 3
Llama 3
Ministral
Phi-4
4B
9B
4B
12B
3B
8B
3B
14B
mini
14B
Mean
KUQ
97.2
98.1
96.8
98.4
96.5
97.8
97.9
98.6
97.3
98.0
97.7
SQuAD 2.0
98.0
98.5
96.2
97.6
95.8
97.1
97.4
98.2
98.3
98.7
97.6
MuSiQue
97.6
98.3
97.1
97.9
96.0
97.3
97.7
98.0
97.5
98.2
97.6
Mean
97.6
98.3
96.7
98.0
96.1
97.4
97.7
98.3
97.7
98.3
97.6
Appendix
Table 6: Agreement (%) between the token-level decision (first token ∈A vs. C ) and the full-generation behavioural label across 30 model-dataset configurations. Mean agreement: 97.6% .
Construction
Sparsity
Ablation AUROC
Main ( ∣A∣=17 , unambiguous, pooled)
0.051
0.641
(i) Per-model sets ( ∣A∣=8 – 15 )
0.054
0.638
(ii) Expanded with prefix check ( ∣A∣=24 )
0.052
0.644
(iii) Further expanded ( ∣A∣=36 )
0.053
0.649
Appendix
Table 7: Sensitivity of CAC properties to alternative A constructions, averaged across ten models. Sparsity: fraction of all components selected into the CAC. Ablation AUROC: AUROC of Δ(x) after full ablation of c -components.
Qwen 3.5
Gemma 3
Llama 3
Ministral
Phi-4
4B
9B
4B
12B
3B
8B
3B
14B
mini
14B
AUROC
KUQ
.708
.678
.742
.906
.733
.763
.779
.874
.858
.845
SQuAD 2.0
.714
.872
.689
.838
.640
.788
.868
.862
.753
.831
MuSiQue
.861
.913
.820
.889
.798
.875
.806
.888
.829
.900
Mean
.761
.821
.750
.878
.723
.809
.818
.875
.813
.859
Acc
KUQ
.644
.660
.672
.821
.666
.684
.710
.778
.765
.745
Appendix
Table 8: Full results for Δ(x) analysis across all 30 model-dataset configurations.
Component counts
C:A ratio ρθ
Model
Screened
a
c
Irrel.
Irrel.%
Sparsity
MLP%
ρ60%
ρ70%
ρ80%
Qwen 4B
81
17
31
33
41
.088
29
6.6 ×
2.4 ×
2.1 ×
Qwen 9B
80
22
20
38
48
.077
14
1.5 ×
1.4 ×
2.0 ×
Gemma 4B
77
15
21
41
53
.118
17
1.8 ×
1.2 ×
1.1 ×
Gemma 12B
72
11
5
56
78
.020
12
1.4 ×
1.3 ×
1.1 ×
Llama 3B
81
20
33
28
35
.076
26
∞
7.3 ×
3.4 ×
Appendix
Table 9: CAC localisation summary across ten models. Sparsity: fraction of all model components ( a + c ) selected into the CAC. MLP%: fraction of CAC components that are MLP sublayers. ρθ : C-to-A count ratio at normalised depth θ —the number of c -components seen per a -component seen up to that depth; ρθ>1 means c -components predominate in layers seen so far (bold); ∞ means no a -components have appeared yet.
Qwen 3.5
Gemma 3
Llama 3
Ministral
Phi-4
4B
9B
4B
12B
3B
8B
3B
14B
mini
14B
Mean
Overall
.882
.908
.854
.922
.862
.890
.878
.866
.882
.898
.884
Answerable
.854
.882
.832
.902
.840
.868
.856
.846
.860
.876
.862
Unanswerable
.910
.934
.876
.942
.884
.912
.900
.886
.904
.920
.907
Appendix
Table 10: Sign agreement between ΔCAC(x) and Δ(x) on the test split.
Gemma 12B
Llama 8B
Gemma 4B
a
c
Irr.%
J(\textsca)
J(\textscc)
J
a
c
Irr.%
J(\textsca)
J(\textscc)
J
a
c
Irr.%
J(\textsca)
J(\textscc)
J
K
11
5
78
1.00
1.00
1.00
15
19
56
1.00
1.00
1.00
15
21
53
1.00
1.00
1.00
2K
11
5
89
1.00
1.00
1.00
15
19
78
1.00
1.00
1.00
16
21
75
0.94
1.00
0.97
4K
12
5
94
0.92
1.00
0.94
16
20
88
0.94
0.95
0.94
17
22
87
0.88
0.95
0.92
Appendix
Table 11: Screening threshold sensitivity. K=72 for Gemma 12B; K=77 for Llama 8B and Gemma 4B. a / c : component counts. Irr.%: fraction of screened candidates rejected by causal gating. J(\textsca) , J(\textscc) , J : per-type and full-CAC Jaccard against the K baseline.
Gemma 12B
Llama 8B
Gemma 4B
λ
a
c
J
Dep.
pˉ
a
c
J
Dep.
pˉ
a
c
J
Dep.
pˉ
0.5λ0
9
4
0.83
+.11
0.58
13
17
0.86
+.12
0.60
13
18
0.85
+.03
0.58
0.75λ0
10
5
0.93
+.12
0.62
14
18
0.93
+.13
0.64
14
20
0.93
+.04
0.63
λ0
11
5
1.00
+.13
—
15
19
1.00
+.14
—
15
21
1.00
+.02
—
1.5λ0
11
6
0.94
+.12
0.61
15
20
0.95
+.13
0.63
16
22
0.94
+.04
0.62
2λ0
12
6
0.88
+.12
0.57
16
21
0.90
+.13
0.59
17
23
0.84
+.05
0.57
Appendix
Table 12: Regularisation sensitivity. λ0=E[∣Δ(x)∣]/(BK) , where B is the batch size. a / c : component counts. J : full-CAC Jaccard against λ0 . Dep.: depth gap ℓˉ\textsca/L−ℓˉ\textscc/L . pˉ : mean cross-seed identification rate of components entering or leaving relative to λ0 .
Ablate a -components
Ablate c -components
A → C
C → A
C → A
A → C
Qwen 4B
6.6
9.5
0.0
45.9
Qwen 9B
9.8
5.7
12.4
12.8
Gemma 4B
19.7
17.9
12.5
7.3
Gemma 12B
1.9
3.9
2.7
3.3
Llama 3B
52.9
2.3
0.1
59.7
Appendix
Table 13: Directional sign flips (% of all 1,500 held-out instances) under full ablation (A → C = abstain flips to commit; C → A = commit flips to abstain).
FA rate (moderate α )
FA rate (extreme α )
mean Δ(x)
Model
1.0
1.5
2.0
2.5
3.0
10
100
1000
α=1
α=2
α=3
Llama 3B
.416
.363
.409
.413
.344
.068
.003
.000
−1.89
−3.24
−0.69
Llama 8B
.475
.415
.408
.409
.417
.001
.000
.000
−3.82
−0.92
−0.64
Qwen 4B
.309
.240
.269
.227
.088
.000
.000
.000
−1.36
−1.32
+1.56
Qwen 9B
.480
.440
.415
.383
.373
.005
.000
.000
−2.06
−1.79
−1.53
Ministral 3B
.320
.237
.224
.231
.047
.140
.001
.001
−1.74
−0.90
+0.90
Appendix
Table 14: FA rate and mean Δ(x) under c -component amplification on FA instances.
FC rate (moderate α )
FC rate (extreme α )
mean Δ(x)
Model
1.0
1.5
2.0
2.5
3.0
10
100
1000
α=1
α=2
α=3
Llama 3B
.337
.311
.280
.237
.228
.316
.331
.332
+2.86
+1.92
+1.05
Llama 8B
.224
.213
.199
.187
.179
.197
.224
.224
+3.55
+2.97
+2.27
Qwen 4B
.380
.311
.247
.172
.147
.368
.380
.379
+2.08
+0.66
−0.37
Qwen 9B
.127
.100
.083
.071
.067
.127
.127
.127
+1.48
+1.05
+1.05
Ministral 3B
.252
.232
.204
.199
.189
.189
.231
.231
+1.99
+1.27
+0.98
Appendix
Table 15: FC rate and mean Δ(x) under a -component amplification on FC instances.
Qwen 3.5
Gemma 3
Llama 3
Ministral
Phi-4
4B
9B
4B
12B
3B
8B
3B
14B
mini
14B
KUQ
Zero-Threshold
.618
.646
.656
.818
.662
.686
.722
.778
.748
.740
Non-CAC Random
.664
.742
.706
.762
.612
.772
.694
.714
.716
.774
INSIDE
.600
.564
.628
.828
.626
.658
.562
.594
.526
.550
Multi-LLM
.730
.672
.815
.854
.640
.800
.705
.794
.691
.727
Appendix
Table 16: Overall accuracy across 10 models and 3 held-out datasets. Bold : best in column for each dataset section; underline : second best.
Qwen 3.5
Gemma 3
Llama 3
Ministral
Phi-4
4B
9B
4B
12B
3B
8B
3B
14B
mini
14B
Precision
KUQ
Zero-Threshold
.595
.693
.610
.761
.651
.670
.691
.723
.754
.752
CAC-Informed MLP (Ours)
.893
.891
.915
.931
.898
.939
.892
.922
.939
.959
SQuAD 2.0
Zero-Threshold
.588
.845
.561
.674
.534
.677
.722
.628
.604
.700
Appendix
Table 17: Commitment-level precision, recall, and F1, and mean false abstention rate, across 10 models and 3 datasets. Bold : best in column. Mean false abstention rate: 0.127 (CAC-Informed MLP) vs. 0.320 (Zero-Threshold).
Qwen 3.5
Gemma 3
Llama 3
Ministral
Phi-4
4B
9B
4B
12B
3B
8B
3B
14B
mini
14B
Mean
KUQ
Zero-Threshold
.618
.646
.656
.818
.662
.686
.722
.778
.748
.740
.707
CAC LR
.878
.872
.884
.914
.856
.854
.908
.910
.888
.910
.887
CAC MLP
.912
.904
.928
.946
.908
.930
.922
.936
.902
.944
.923
SQuAD 2.0
Appendix
Table 18: Accuracy across 10 models and 3 in-distribution datasets. Bold : best in row for each dataset section. Mean : averaged over all 3 datasets.
Table 24
Qwen 3.5
Gemma 3
Llama 3
Ministral
Phi-4
4B
9B
35B
4B
12B
27B
3B
8B
3B
14B
mini
14B
Total
Δ(x) evaluation
6.2
13.4
20.6
6.7
18.1
39.7
4.3
11.0
4.1
16.9
4.6
18.2
163.8
CAC localisation
16.2
35.8
50.4
11.3
38.6
77.9
9.1
28.7
8.4
31.2
9.3
38.9
355.8
Ablation
3.8
9.7
12.1
4.1
11.6
27.3
3.2
8.4
2.9
11.8
3.1
12.2
110.2
Amplification
4.9
11.8
14.3
5.3
14.7
34.6
3.7
10.6
3.8
14.9
4.2
15.7
138.5
Policy training
1.4
2.8
3.6
1.6
3.7
8.6
1.1
2.6
0.9
3.5
0.8
3.9
34.5
Appendix
Table 21: Estimated wall-clock time (minutes) per pipeline stage on a single NVIDIA H100 80GB (1,500 data instances, 10 seeds for CAC localisation).
Department of Mechanical and Aerospace University of California Irvine Irvine, CA 92617-4322, USA · T-5 Los Alamos National Laboratory Los Alamos, NM 88220, USA · CAI-4 Los Alamos National Laboratory Los Alamos, NM 88220, USA