Language models (LMs) often hallucinate by committing to confident answers rather than abstaining, even when they do not have enough information to answer reliably. A large body of existing work mitigates hallucination through detection or abstention mechanisms, but leaves open how models internally arrive at the decision to commit or abstain in the first place. We study this decision through mechanistic analysis, framing hallucination as unsupported commitment: the model commits despite exhibiting signals of unanswerability. Using causal gating, we identify a Commit-Abstain Circuit (CAC), a sparse, causally localised subset of attention heads and MLP sublayers underlying this decision. Across ten LMs (3B-14B) from five families and three benchmarks, the CAC exhibits a recurring accumulate-yet-undercorrect pattern: commitment-promoting components build up commitment in earlier layers, while abstention-promoting components act later as corrective signals that are often insufficient to overturn the accumulated commitment. Building on this finding, a lightweight policy trained on CAC activations improves decision accuracy by 12.2 points over the model's intrinsic commit-abstain margin, reduces false abstentions by 2.5 times, transfers to unseen benchmarks, and extends to larger models (27B-35B). The CAC is both diagnostic, clarifying how models overcommit, and practical, enabling improved abstention decisions.
Figures & tables
Figure 1: Unsupported commitment in Qwen 3.5 4B on an unanswerable KUQ question [ 27 ] . Left: Per-component contributions to the commit-abstain margin Δ(x) (Definition 1 ). Late-layer abstention signals fail to overturn earlier commitment ( Δ(x)=+3.25 ), leading to commitment. Right: A lightweight CAC-Informed MLP correctly triggers abstention.
Abstention set A
Commitment set C
“Don’t know”, “Insufficient evidence”, “Unsure”, “Unanswerable”, “Ambiguous”, “Beyond my knowledge”
All remaining, e.g. those initiating: “According to…”, “The answer is…”, “Exactly”, “Yes”, “No”, “In fact”
Table 1: Token sets used to define the commit-abstain margin Δ(x) .
Figure 3Figure 4Figure 5
Qwen 3.5
Gemma 3
Llama 3
Ministral
Phi-4
4B
9B
4B
12B
3B
8B
3B
14B
mini
14B
p
Mean AUROC
Zero-Threshold
.725
.798
.722
.854
.689
.784
.790
.848
.780
.830
6.0×10−6
Non-CAC Random
.691
.764
.741
.808
.648
.815
.732
.789
.756
.822
8.4×10−6
INSIDE
.572
.449
.586
.658
.578
.483
.406
.493
.415
.492
3.7×10−9
Multi-LLM
.688
.757
.683
.728
.538
.705
.596
.654
.603
.703
1.9×10−9
Semantic Entropy
.567
.463
.548
.600
.564
.521
.423
.564
.509
.529
1.3×10−8
Table 2: Mean AUROC and accuracy across three datasets for each model. Bold indicates the best result; underline indicates the second best. p -values are from two-sided Wilcoxon signed-rank tests over 30 model-dataset pairs. All differences are significant ( p<0.05 ), with most at p<10−5 .
Qwen 3.5
Gemma 3
Llama 3
Ministral
Phi-4
4B
9B
35B †
4B
12B
27B †
3B
8B
3B
14B
mini
14B
HotpotQA
Zero-Threshold
.887
.837
.867
.898
.922
.924
.840
.820
.851
.892
.859
.842
HaMI
.917
.938
.925
.894
.915
.892
.829
.924
.865
.889
.856
.937
CAC-Informed MLP (Ours)
.930
.949
.928
.911
.948
.950
.859
.913
.895
.894
.875
.939
SelfAware
Table 3: Decision accuracy on two unseen datasets (HotpotQA, SelfAware) and two unseen larger models ( † ). Our method achieves higher accuracy in all 24 configurations ( p=1.8×10−5 ).
Appendix figures & tables18 assets
Supplementary material from the paper’s appendix.
Appendix
Dataset
Source
Reasoning
Unanswerability
License
KUQ [ 27 ]
Parametric
Open-domain QA
Answers genuinely do not exist (unsolved problems, future events); tests limits of human knowledge.
MIT
SQuAD 2.0 [ 78 ]
In-context
Reading comprehension
Questions plausible but unsupported by the passage; tests resistance to surface-level cues.
CC BY-SA 4.0
MuSiQue [ 79 ]
In-context
Multi-hop
One supporting document removed ( missing evidence ); tests detection of a silent gap in the reasoning chain.
CC BY 4.0
SelfAware † [ 92 ]
Parametric
Self-knowledge
No definitive answer exists; tests limits of the model’s own knowledge, not of the world.
CC BY-SA 4.0
HotpotQA † [ 91 ]
In-context
Multi-hop
Supporting passages replaced by distractors ( misleading evidence ); tests resistance to plausible-but-wrong context.
CC BY-SA 4.0
Appendix
Table 4: Dataset summary. Source : whether the model must rely on parametric knowledge or a provided context passage. ( † ) held-out OOD evaluation sets only.
Figure 9: Annotation interface shown to Amazon Mechanical Turk workers for classifying candidate tokens as epistemic deferral or substantive answer.
I
“I don’t know”, “I cannot determine”
Sorry
“Sorry, I don’t have information…”
Unfortunately
“Unfortunately, I don’t have enough…”
Beyond
“Beyond my knowledge…”
Insufficient
“Insufficient information to answer…”
Impossible
“Impossible to determine without context”
Unanswerable
“Unanswerable based on context”
Without
“Without more context, I cannot…”
Unclear
“Unclear from the given context…”
Ambiguous
“Ambiguous; the premise is unclear”
Unknown
“Unknown; context does not specify…”
Unsure
“Unsure; evidence is inconclusive”
Uncertain
“Uncertain; lacks a definitive answer”
Unknowable
“Unknowable from the text provided”
Appendix
Table 5: Abstention token set A ( 17 tokens). The commitment set C contains all tokens not in A .
Qwen 3.5
Gemma 3
Llama 3
Ministral
Phi-4
4B
9B
4B
12B
3B
8B
3B
14B
mini
14B
Mean
KUQ
97.2
98.1
96.8
98.4
96.5
97.8
97.9
98.6
97.3
98.0
97.7
SQuAD 2.0
98.0
98.5
96.2
97.6
95.8
97.1
97.4
98.2
98.3
98.7
97.6
MuSiQue
97.6
98.3
97.1
97.9
96.0
97.3
97.7
98.0
97.5
98.2
97.6
Mean
97.6
98.3
96.7
98.0
96.1
97.4
97.7
98.3
97.7
98.3
97.6
Appendix
Table 6: Agreement (%) between the token-level decision (first token ∈A vs. C ) and the full-generation behavioural label across 30 model-dataset configurations. Mean agreement: 97.6% .
Construction
Sparsity
Ablation AUROC
Main ( ∣A∣=17 , unambiguous, pooled)
0.051
0.641
(i) Per-model sets ( ∣A∣=8 – 15 )
0.054
0.638
(ii) Expanded with prefix check ( ∣A∣=24 )
0.052
0.644
(iii) Further expanded ( ∣A∣=36 )
0.053
0.649
Appendix
Table 7: Sensitivity of CAC properties to alternative A constructions, averaged across ten models. Sparsity: fraction of all components selected into the CAC. Ablation AUROC: AUROC of Δ(x) after full ablation of c -components.
Qwen 3.5
Gemma 3
Llama 3
Ministral
Phi-4
4B
9B
4B
12B
3B
8B
3B
14B
mini
14B
AUROC
KUQ
.708
.678
.742
.906
.733
.763
.779
.874
.858
.845
SQuAD 2.0
.714
.872
.689
.838
.640
.788
.868
.862
.753
.831
MuSiQue
.861
.913
.820
.889
.798
.875
.806
.888
.829
.900
Mean
.761
.821
.750
.878
.723
.809
.818
.875
.813
.859
Acc
KUQ
.644
.660
.672
.821
.666
.684
.710
.778
.765
.745
Appendix
Table 8: Full results for Δ(x) analysis across all 30 model-dataset configurations.
Component counts
C:A ratio ρθ
Model
Screened
a
c
Irrel.
Irrel.%
Sparsity
MLP%
ρ60%
ρ70%
ρ80%
Qwen 4B
81
17
31
33
41
.088
29
6.6 ×
2.4 ×
2.1 ×
Qwen 9B
80
22
20
38
48
.077
14
1.5 ×
1.4 ×
2.0 ×
Gemma 4B
77
15
21
41
53
.118
17
1.8 ×
1.2 ×
1.1 ×
Gemma 12B
72
11
5
56
78
.020
12
1.4 ×
1.3 ×
1.1 ×
Llama 3B
81
20
33
28
35
.076
26
∞
7.3 ×
3.4 ×
Appendix
Table 9: CAC localisation summary across ten models. Sparsity: fraction of all model components ( a + c ) selected into the CAC. MLP%: fraction of CAC components that are MLP sublayers. ρθ : C-to-A count ratio at normalised depth θ —the number of c -components seen per a -component seen up to that depth; ρθ>1 means c -components predominate in layers seen so far (bold); ∞ means no a -components have appeared yet.
Qwen 3.5
Gemma 3
Llama 3
Ministral
Phi-4
4B
9B
4B
12B
3B
8B
3B
14B
mini
14B
Mean
Overall
.882
.908
.854
.922
.862
.890
.878
.866
.882
.898
.884
Answerable
.854
.882
.832
.902
.840
.868
.856
.846
.860
.876
.862
Unanswerable
.910
.934
.876
.942
.884
.912
.900
.886
.904
.920
.907
Appendix
Table 10: Sign agreement between ΔCAC(x) and Δ(x) on the test split.
Gemma 12B
Llama 8B
Gemma 4B
a
c
Irr.%
J(\textsca)
J(\textscc)
J
a
c
Irr.%
J(\textsca)
J(\textscc)
J
a
c
Irr.%
J(\textsca)
J(\textscc)
J
K
11
5
78
1.00
1.00
1.00
15
19
56
1.00
1.00
1.00
15
21
53
1.00
1.00
1.00
2K
11
5
89
1.00
1.00
1.00
15
19
78
1.00
1.00
1.00
16
21
75
0.94
1.00
0.97
4K
12
5
94
0.92
1.00
0.94
16
20
88
0.94
0.95
0.94
17
22
87
0.88
0.95
0.92
Appendix
Table 11: Screening threshold sensitivity. K=72 for Gemma 12B; K=77 for Llama 8B and Gemma 4B. a / c : component counts. Irr.%: fraction of screened candidates rejected by causal gating. J(\textsca) , J(\textscc) , J : per-type and full-CAC Jaccard against the K baseline.
Gemma 12B
Llama 8B
Gemma 4B
λ
a
c
J
Dep.
pˉ
a
c
J
Dep.
pˉ
a
c
J
Dep.
pˉ
0.5λ0
9
4
0.83
+.11
0.58
13
17
0.86
+.12
0.60
13
18
0.85
+.03
0.58
0.75λ0
10
5
0.93
+.12
0.62
14
18
0.93
+.13
0.64
14
20
0.93
+.04
0.63
λ0
11
5
1.00
+.13
—
15
19
1.00
+.14
—
15
21
1.00
+.02
—
1.5λ0
11
6
0.94
+.12
0.61
15
20
0.95
+.13
0.63
16
22
0.94
+.04
0.62
2λ0
12
6
0.88
+.12
0.57
16
21
0.90
+.13
0.59
17
23
0.84
+.05
0.57
Appendix
Table 12: Regularisation sensitivity. λ0=E[∣Δ(x)∣]/(BK) , where B is the batch size. a / c : component counts. J : full-CAC Jaccard against λ0 . Dep.: depth gap ℓˉ\textsca/L−ℓˉ\textscc/L . pˉ : mean cross-seed identification rate of components entering or leaving relative to λ0 .
Ablate a -components
Ablate c -components
A → C
C → A
C → A
A → C
Qwen 4B
6.6
9.5
0.0
45.9
Qwen 9B
9.8
5.7
12.4
12.8
Gemma 4B
19.7
17.9
12.5
7.3
Gemma 12B
1.9
3.9
2.7
3.3
Llama 3B
52.9
2.3
0.1
59.7
Appendix
Table 13: Directional sign flips (% of all 1,500 held-out instances) under full ablation (A → C = abstain flips to commit; C → A = commit flips to abstain).
FA rate (moderate α )
FA rate (extreme α )
mean Δ(x)
Model
1.0
1.5
2.0
2.5
3.0
10
100
1000
α=1
α=2
α=3
Llama 3B
.416
.363
.409
.413
.344
.068
.003
.000
−1.89
−3.24
−0.69
Llama 8B
.475
.415
.408
.409
.417
.001
.000
.000
−3.82
−0.92
−0.64
Qwen 4B
.309
.240
.269
.227
.088
.000
.000
.000
−1.36
−1.32
+1.56
Qwen 9B
.480
.440
.415
.383
.373
.005
.000
.000
−2.06
−1.79
−1.53
Ministral 3B
.320
.237
.224
.231
.047
.140
.001
.001
−1.74
−0.90
+0.90
Appendix
Table 14: FA rate and mean Δ(x) under c -component amplification on FA instances.
FC rate (moderate α )
FC rate (extreme α )
mean Δ(x)
Model
1.0
1.5
2.0
2.5
3.0
10
100
1000
α=1
α=2
α=3
Llama 3B
.337
.311
.280
.237
.228
.316
.331
.332
+2.86
+1.92
+1.05
Llama 8B
.224
.213
.199
.187
.179
.197
.224
.224
+3.55
+2.97
+2.27
Qwen 4B
.380
.311
.247
.172
.147
.368
.380
.379
+2.08
+0.66
−0.37
Qwen 9B
.127
.100
.083
.071
.067
.127
.127
.127
+1.48
+1.05
+1.05
Ministral 3B
.252
.232
.204
.199
.189
.189
.231
.231
+1.99
+1.27
+0.98
Appendix
Table 15: FC rate and mean Δ(x) under a -component amplification on FC instances.
Qwen 3.5
Gemma 3
Llama 3
Ministral
Phi-4
4B
9B
4B
12B
3B
8B
3B
14B
mini
14B
KUQ
Zero-Threshold
.618
.646
.656
.818
.662
.686
.722
.778
.748
.740
Non-CAC Random
.664
.742
.706
.762
.612
.772
.694
.714
.716
.774
INSIDE
.600
.564
.628
.828
.626
.658
.562
.594
.526
.550
Multi-LLM
.730
.672
.815
.854
.640
.800
.705
.794
.691
.727
Appendix
Table 16: Overall accuracy across 10 models and 3 held-out datasets. Bold : best in column for each dataset section; underline : second best.
Qwen 3.5
Gemma 3
Llama 3
Ministral
Phi-4
4B
9B
4B
12B
3B
8B
3B
14B
mini
14B
Precision
KUQ
Zero-Threshold
.595
.693
.610
.761
.651
.670
.691
.723
.754
.752
CAC-Informed MLP (Ours)
.893
.891
.915
.931
.898
.939
.892
.922
.939
.959
SQuAD 2.0
Zero-Threshold
.588
.845
.561
.674
.534
.677
.722
.628
.604
.700
Appendix
Table 17: Commitment-level precision, recall, and F1, and mean false abstention rate, across 10 models and 3 datasets. Bold : best in column. Mean false abstention rate: 0.127 (CAC-Informed MLP) vs. 0.320 (Zero-Threshold).
Qwen 3.5
Gemma 3
Llama 3
Ministral
Phi-4
4B
9B
4B
12B
3B
8B
3B
14B
mini
14B
Mean
KUQ
Zero-Threshold
.618
.646
.656
.818
.662
.686
.722
.778
.748
.740
.707
CAC LR
.878
.872
.884
.914
.856
.854
.908
.910
.888
.910
.887
CAC MLP
.912
.904
.928
.946
.908
.930
.922
.936
.902
.944
.923
SQuAD 2.0
Appendix
Table 18: Accuracy across 10 models and 3 in-distribution datasets. Bold : best in row for each dataset section. Mean : averaged over all 3 datasets.
Table 24
Qwen 3.5
Gemma 3
Llama 3
Ministral
Phi-4
4B
9B
35B
4B
12B
27B
3B
8B
3B
14B
mini
14B
Total
Δ(x) evaluation
6.2
13.4
20.6
6.7
18.1
39.7
4.3
11.0
4.1
16.9
4.6
18.2
163.8
CAC localisation
16.2
35.8
50.4
11.3
38.6
77.9
9.1
28.7
8.4
31.2
9.3
38.9
355.8
Ablation
3.8
9.7
12.1
4.1
11.6
27.3
3.2
8.4
2.9
11.8
3.1
12.2
110.2
Amplification
4.9
11.8
14.3
5.3
14.7
34.6
3.7
10.6
3.8
14.9
4.2
15.7
138.5
Policy training
1.4
2.8
3.6
1.6
3.7
8.6
1.1
2.6
0.9
3.5
0.8
3.9
34.5
Appendix
Table 21: Estimated wall-clock time (minutes) per pipeline stage on a single NVIDIA H100 80GB (1,500 data instances, 10 seeds for CAC localisation).
Modern language models fail a fundamental requirement of trustworthy intelligence: knowing when not to answer. Despite achieving impressive accuracy on benchmarks, these models produce confident hallucinations, even when wrong answers carry catastrophic consequences. Our evaluations on GSM8K, MedQA and GPQA show frontier models almost never abstain despite explicit warnings of severe penalties, suggesting that prompts cannot override training that rewards any answer over no answer. As a remedy, we propose Reinforced Hesitation (RH): a modification to Reinforcement Learning from Verifiable Rewards (RLVR) to use ternary rewards (+1 correct, 0 abstention, -λ error) instead of binary. Controlled experiments on logic puzzles reveal that varying λ produces distinct models along a Pareto frontier, where each training penalty yields the optimal model for its corresponding risk regime: low penalties produce aggressive answerers, high penalties conservative abstainers. The same frontier holds on MATH Levels 4--5 and on medical QA, where it transfers to an unseen dataset. We then introduce two inference strategies that exploit trained abstention as a coordination signal: cascading routes queries through models with decreasing risk tolerance, while self-cascading re-queries the same model on abstention. Both outperform majority voting with lower computational cost. These results establish abstention as a first-class training objective that transforms ``I don't know'' from failure into a coordination signal, enabling models to earn trust through calibrated honesty about their limits.
Mohamad Amin Mohamadi, Tianhao Wang, Zhiyuan Li
Toyota Technological Institute at Chicago · University of California, San Diego
When language models lack relevant knowledge for a given query, they frequently generate plausible responses that can be hallucinations, rather than admitting being agnostic about the answer. Retraining models to reward admitting ignorance can lead to overly conservative behaviors and poor generalization due to scarce evaluation benchmarks. We propose a post hoc framework, Conformal Abstention (CA), adapted from conformal prediction (CP) to determine whether to abstain from answering a query. CA provides finite-sample guarantees on both the probability of participation (i.e., not abstaining) and the probability that the generated response is correct. Importantly, the abstention decision relies on prediction confidence rather than the non-conformity scores used in CP, which are intractable for open-ended generation. To better align prediction confidence with the model's ignorance, we introduce a calibration strategy using representation geometry within the model to measure knowledge involvement in shaping the response. Experiments demonstrate that we improve selective answering significantly with 75 percent conditional correctness.
Rui Xu, Yi Chen, Sihong Xie +1
Information Hub, AI Thrust Hong Kong University of Science and Technology (Guangzhou) Guangzhou, Guangdong, China
We introduce CAROL (Chain-based Adaptive Reconfiguration Over Lattices), a probabilistic framework for test-time hallucination reduction in large language models. Rather than relying on token-level uncertainty, CAROL defines a semantic uncertainty measure based on the consistency between generated responses and a trusted context, inducing a string-submodular objective over a lattice of textual sequences. This formulation enables hallucination mitigation to be cast as a Markov chain accept-reject process with provable convergence and near-optimality guarantees, allowing the model to iteratively refine outputs toward semantic consistency. By operating at the level of meaning, CAROL unifies hallucination detection and mitigation within a single framework. Empirical results on question answering and multi-agent reasoning benchmarks show that CAROL significantly reduces hallucinations and improves reliability and interpretability compared to likelihood-based and retrieval-augmented baselines, while maintaining competitive computational efficiency.
Joan Vendrell Gallart, Solmaz Kia, Russell Bent +1
Department of Mechanical and Aerospace University of California Irvine Irvine, CA 92617-4322, USA · T-5 Los Alamos National Laboratory Los Alamos, NM 88220, USA · CAI-4 Los Alamos National Laboratory Los Alamos, NM 88220, USA