Modern language models fail a fundamental requirement of trustworthy intelligence: knowing when not to answer. Despite achieving impressive accuracy on benchmarks, these models produce confident hallucinations, even when wrong answers carry catastrophic consequences. Our evaluations on GSM8K, MedQA and GPQA show frontier models almost never abstain despite explicit warnings of severe penalties, suggesting that prompts cannot override training that rewards any answer over no answer. As a remedy, we propose Reinforced Hesitation (RH): a modification to Reinforcement Learning from Verifiable Rewards (RLVR) to use ternary rewards (+1 correct, 0 abstention, -λ error) instead of binary. Controlled experiments on logic puzzles reveal that varying λ produces distinct models along a Pareto frontier, where each training penalty yields the optimal model for its corresponding risk regime: low penalties produce aggressive answerers, high penalties conservative abstainers. The same frontier holds on MATH Levels 4--5 and on medical QA, where it transfers to an unseen dataset. We then introduce two inference strategies that exploit trained abstention as a coordination signal: cascading routes queries through models with decreasing risk tolerance, while self-cascading re-queries the same model on abstention. Both outperform majority voting with lower computational cost. These results establish abstention as a first-class training objective that transforms ``I don't know'' from failure into a coordination signal, enabling models to earn trust through calibrated honesty about their limits.
Figures & tables
Figure 1 : Reinforced Hesitation creates an ordered risk frontier on Knights & Knaves, reducing error rates through calibrated abstentions and enabling adaptive inference. All panels in this summary figure use Knights & Knaves. Right: Cross-penalty evaluation reveals risk specialization across our model family: each model achieves superior performance under specific evaluation penalties, with optimal models (orange) clustering near the diagonal where training and evaluation contexts align. This demonstrates that each training penalty produces a useful operating point for a corresponding risk regime. Top left: Cascading through models with decreasing risk aversion ( λ=10→5→2→1→0 ) achieves efficient triage where each specialist handles problems matching its confidence regime. Bottom left: Models trained with different penalties form a frontier where higher λ achieves both lower error rates and higher conditional accuracy.
Figure 2 : Penalty sensitivity of frontier models on GSM8K. Left: Expected reward r(λ)=p(correct)−λp(wrong) for λ∈{1,5,25,100} ; red dashed line marks r=0 baseline. Middle: Frontier models rarely choose to abstain, even when faced with penalties of magnitude 100. Right: Despite the high penalty values, the rate of wrong answers remains high across various models.
Figure 3 : Knights & Knaves behavioral decomposition by difficulty. Solid lines show easy logic puzzles (66% of the Knights & Knaves dataset), dashed lines show hard puzzles (33%). Left: Models with λ>0 learn calibrated abstention: 5-15% on easy versus 60-95% on hard problems, while λ=0 abstains near 0% regardless. The λ=10 transient spike to 97% on easy problems (step 80) followed by recovery indicates behavioral recalibration. Middle: Wrong rates rapidly suppress to < 2% for all λ≥1 while λ=0 maintains 10-20% errors. Right: Correct rates reveal the coverage-safety tradeoff, with λ=10 achieving 90% accuracy when it does answer.
Figure 4 : MATH validation metrics. Fixed-penalty RH specialists show the intended trade-off directly: increasing λ lowers wrong-answer rate and raises conditional accuracy, while abstention absorbs coverage. Full outcomes, risk-frontier, difficulty, and format diagnostics are in Appendix D.3 .
Figure 5 : Knights & Knaves validation performance reveals learned selectivity. As training penalty λ increases, models trade coverage for safety on held-out logic puzzles: correct rates decrease while abstention rates rise, but wrong rates collapse dramatically. The most informative insight is conditional accuracy (rightmost panel) jumping from 84% to > 99%, proving models learn to abstain precisely on problems where they would likely make mistakes.
Figure 6 : Knights & Knaves cascade is query-efficient through behavioral complementarity. Dashed lines represent majority voting applied on models trained with different penalties, solid lines represent cascading. Left: Cascade ( ★ ) achieves 88.1% accuracy with only 2.2 average queries, outperforming individual Knights & Knaves specialists. Middle: IDK rates collapse from 33% to < 1% through cascading. Right: Wrong rates remain controlled at 11.5%, competitive with the baseline while having higher coverage.
Figure 7 : Knights & Knaves self-cascading converts abstentions to answers through nondeterminism. Left: Correct rate increases with budget for models with λ≥1 , with λ=1 showing steepest gains (77.5% → 92.5%). The λ=0 baseline remains flat as it never abstains. Middle: IDK rates decay with budget. Right: Wrong rates increase as abstentions convert to answers while mistakes remain bounded.
Appendix figures & tables27 assets
Supplementary material from the paper’s appendix.
Appendix
Method
Uncertainty signal
Training objective
Relation to RH
RH (ours)
Explicit answer/IDK action
Ternary utility (+1,0,−λ)
Learns fixed-risk specialists and uses abstention as an inference-routing signal.
TruthRL ( Wei et al., 2026 )
Explicit abstention
Fixed ternary abstention reward
Closest to the λ=1 member of RH in verifiable-answer settings.
RLCR ( Damani et al., 2026 )
Verbalized confidence
Correctness plus Brier-style calibration
Direct confidence-threshold baseline; we threshold at λ/(1+λ) .
BCRL ( Wu et al., 2026 )
Verbalized confidence, thresholded post hoc at a user-specified risk level t
Proper scoring rules; also a prompted- t ternary reward (+1,0,−t/(1−t))
Adjacent single-model approach to user-specified risk; its prompted- t variant did not adapt to t .
Rewarding Doubt ( Bani-Harouni et al., 2026 )
Verbalized integer confidence (0 to 10)
Logarithmic scoring rule on confidence, with answers held fixed
Complementary direction focused on calibrated confidence expression.
Appendix
Table 1 : Closest training-time abstention and confidence methods. RH is closest to fixed ternary abstention training, but the main empirical baseline in this paper is RLCR-style verbalized confidence ( Damani et al., 2026 ) , which we threshold at the Bayes cutoff λ/(1+λ) to obtain abstention.
Suite
Model and data
Reward variants
Knights & Knaves
Qwen/Qwen3-1.7B ( Qwen Team, 2025 ) ; 80,000 training puzzles and 10,000 test puzzles.
Fixed RH penalties λ∈{0,1,2,5,10,20} .
MATH
Qwen3.5-4B; MATH Levels 4–5 with 3,994 training problems and 2,538 validation problems.
Fixed RH penalties λ∈{0,1,2,5,10} , prompt-conditioned dynamic penalty over the same set, expected-penalty ablation at λ=3.6 , and an RLCR baseline ( Damani et al., 2026 ) .
Appendix
Table 2 : Training suites. The two suites share the same answer/abstain decision structure while testing RH on different reasoning domains and model scales.
Setting
Value
Shared settings
Training pipeline
verl ( Sheng et al., 2025 ) with SGLang rollouts ( Zheng et al., 2024 ) ; MATH uses the same verl/DrGRPO stack.
Rollout sampling
Temperature 1.0 with nucleus sampling (top-p=1.0), 8 rollouts per prompt.
Validation decoding
Deterministic decoding, temperature 0, n=1 .
Response limit
4096 tokens.
Knights & Knaves
Appendix
Table 3 : Core training configuration. Shared settings are listed first, followed by suite-specific optimization and evaluation details.
Figure 8 : MedQA results parallel to GSM8K (Figure 2 ). Left: Expected reward shows all models fall below zero for λ≥5 . Middle: The hesitation curve is flat across all models and penalty conditions. Right: Wrong answer rates are largely unchanged (6-36%) under prompt-level penalties.
Figure 9 : GPQA results showing domain-dependent calibration. Left: Expected rewards are lower due to the challenging baseline accuracy. Middle: Several models show penalty-sensitive abstention, with GPT-4o reaching 20.76% at λ=100 . Right: Wrong rates decrease as models abstain more on difficult graduate-level questions.
Figure 10 : Accuracy across all datasets and penalty conditions. This comprehensive view shows how model accuracy varies with penalty magnitude across GSM8K (top), MedQA (middle), and GPQA (bottom). While most models maintain stable accuracy on GSM8K and MedQA regardless of penalty, GPQA shows more variability, with some models (e.g., GPT-4o) trading accuracy for reduced error rates through strategic abstention at higher penalties.
Figure 11 : Training dynamics across penalty values. Abstention, error, and correct answer rates during training for all penalty values λ∈{0,1,2,5,10,20} . Left: Abstention rates remain near zero for λ≤5 while λ=10 shows a transient increase at step 80 before stabilizing at 20-30%. Middle: Wrong answer rates decrease with higher penalties, reaching 10-15% for λ≥10 . Right: Correct answer rates stabilize between 40-70% depending on penalty value, with λ=0 maintaining highest coverage and λ=20 selecting lower coverage through increased abstention.
Figure 12 : Response generation metrics during training. Left: Mean response length across training steps for each penalty value. Middle: Fraction of responses reaching the 4096 token limit. Right: Format error rates by difficulty (solid lines: easy problems, dashed lines: hard problems). Format-error rates vary during behavioral transitions, with hard problems requiring more formatting control.
λ
Correct
IDK
Wrong
Format err.
Cond. acc.
0
88.48%
0.08%
11.44%
1.59%
88.6%
1
87.61%
6.13%
6.25%
0.21%
93.3%
2
85.78%
8.16%
6.05%
0.11%
93.4%
5
81.81%
12.96%
5.23%
0.00%
94.0%
10
62.85%
34.39%
2.75%
0.04%
95.8%
Appendix
Table 4 : Full single-sample MATH validation outcomes by training penalty. Rates are averaged over 16 validation rollouts per problem at the selected checkpoints. Conditional accuracy is correct divided by non-IDK terminal outcomes under the conservative scoring policy.
Figure 13 : MATH fixed-penalty training dynamics. Across training, increasing λ shifts behavior toward abstention and away from wrong answers. The validation-selected checkpoints used in the main text are chosen by validation RH reward.
Figure 14 : MATH risk-frontier diagnostics. The frontier and cross-penalty matrix show the softer version of the Knights & Knaves pattern: fixed penalties produce ordered operating points, but neighboring penalties can be competitive under the same evaluation cost.
Deployed model
Eval λ=0
Eval λ=1
Eval λ=2
Eval λ=5
Eval λ=10
Specialist λ=0
0.890
0.781
0.671
0.344
-0.202
Specialist λ=1
0.883
0.818
0.753
0.558
0.233
Specialist λ=2
0.868
0.809
0.749
0.571
0.273
Specialist λ=5
0.841
0.785
0.729
0.561
0.282
Specialist λ=10
0.640
0.615
0.591
0.516
0.392
Appendix
Table 5 : Cross-penalty reward matrix on MATH. Entries are mean validation utility Uλ=correct−λ⋅wrong at the selected fixed-specialist checkpoints. They come from the deterministic (temperature 0) validation pass used for checkpoint selection (Table 3 ) and therefore differ slightly from the 16-rollout averages in Table 4 . Bold marks the best fixed-specialist row for each evaluation penalty.
Figure 15 : MATH performance by difficulty split. Higher penalties produce stronger abstention and lower error rates across problem levels, with the largest coverage reduction at λ=10 .
Figure 16 : MATH response-length and format diagnostics. These diagnostics track average response length, clipping, and format behavior across penalty values, helping separate abstention effects from formatting artifacts.
RLCR + threshold
Matching RH specialist
λ
τλ
Correct
Wrong
IDK
Correct
Wrong
IDK
ΔU
0
0.00
88.7
11.3
0.0
88.5
11.4
0.1
-0.3
1
0.50
87.3
7.4
5.3
87.6
6.3
6.1
+1.5
2
0.67
80.9
5.8
13.2
85.8
6.1
8.2
+4.4
5
0.83
80.9
5.8
13.3
81.8
5.2
13.0
+4.0
10
0.91
42.7
2.6
54.7
62.9
2.8
34.4
+18.6
Appendix
Table 6 : Matched comparison between RLCR thresholding and RH specialists ( Damani et al., 2026 ) . Rates are percentages averaged over 16 validation rollouts per problem. ΔU is RH minus RLCR expected utility in percentage points, using Uλ=correct−λ⋅wrong . Format errors and empty answers are counted as wrong.
Figure 18 : RLCR thresholding and RH specialists respond differently to high error penalties ( Damani et al., 2026 ) . RLCR remains competitive when error costs are low, while higher thresholds shift more cases into abstention at large λ . Matching RH specialists retain more coverage while keeping wrong rates low, yielding higher expected utility for nonzero penalties in Table 6 .
Figure 19 : Majority voting provides a compute-scaling baseline. Left: Correct rates show modest improvement with budget: λ=0 gains 3% (85% → 88%), while λ=1 gains 2% (77.5% → 79.5%). Middle: IDK rates track the model’s selected coverage level under aggregation. Right: Wrong rates decrease through consensus filtering.
Figure 20 : MATH cascade decomposition. Dashed lines represent majority voting applied to fixed-penalty specialists, while the solid green line represents the cross-penalty cascade ordered as λ=10→5→2→1→0 . The full cascade reaches 89.97% correct with 1.58 average queries, collapses IDK to 0.07%, and ends at 9.96% wrong.
Figure 21 : MATH self-cascading. Re-querying a single fixed- λ specialist converts some abstentions into answers. Gains are smaller than on Knights & Knaves but preserve the same early-exit interpretation.
Figure 22 : MATH majority-vote baseline. Majority voting is a strong compute-scaling baseline, especially for the λ=0 model, and motivates our framing of cascades as query-efficient rather than uniformly accuracy-dominant.
Cascade order
Correct
IDK
Wrong
Format err.
Avg. queries
10
62.85
34.39
2.75
0.04
1.00
10→5
82.62
12.01
5.37
0.05
1.34
10→5→2
86.69
7.08
6.23
0.07
1.46
10→5→2→1
88.54
4.60
6.86
0.14
1.53
10→5→2→1→0
89.97
0.07
9.96
0.86
1.58
Appendix
Table 7 : MATH cascade depth diagnostics. Rates are percentages. Adding lower-penalty tiers recovers coverage and correct answers at small additional query cost, while the final λ=0 tier also reintroduces more wrong answers.
Method
Budget
Correct
IDK
Wrong
Single λ=0
1
88.48
0.08
11.44
Majority vote λ=0
4
90.20
0.07
9.73
Majority vote λ=0
8
90.80
0.06
9.13
Majority vote λ=0
16
91.05
0.08
8.87
Full cascade
1.58 avg.
89.97
0.07
9.96
Appendix
Table 8 : Majority vote is the strongest raw-accuracy baseline. Rates are percentages. The full cascade is best interpreted as a query-efficient alternative, not as uniformly accuracy-dominant over 16-query majority voting.
Figure 23 : Dynamic penalty versus expected-penalty training. The prompt-conditioned dynamic model tracks the fixed λˉ=3.6 ablation closely, supporting the main-text interpretation that naive dynamic training partly collapses toward expected-penalty behavior.
Figure 24 : Prompt-conditioned behavior during dynamic training. The dynamic model is not completely penalty-blind, but the prompted- λ effect remains much smaller than the span obtained by separately trained specialists.
Example
Correct at λ =0
Correct at λ =10
IDK at λ =10
1
100.0%
28.1%
71.9%
2
100.0%
31.2%
68.8%
3
98.4%
21.9%
78.1%
4
87.5%
0.0%
100.0%
5
85.9%
0.0%
100.0%
Appendix
Table 9 : Examples of selective coverage at high penalty values. These problems show how high- λ specialists prioritize low-error operation by deferring cases that lower-penalty models often answer, illustrating the coverage side of the accuracy-coverage frontier.
MedMCQA (in-domain)
MedQA (transfer)
λ
Threshold
Coverage
Wrong
Cond. acc.
Coverage
Wrong
Cond. acc.
0
n/a
99.98%
49.91%
50.1%
100.0%
47.50%
52.5%
1
50.0%
53.76%
21.87%
59.3%
75.60%
32.78%
56.6%
2
66.7%
21.06%
5.17%
75.5%
26.02%
7.08%
72.8%
5
83.3%
7.25%
0.94%
87.0%
5.61%
0.80%
85.7%
Appendix
Table 10: Medical suite outcomes at step 150. Rates are averaged over 16 sampled rollouts per question (temperature 1.0). Coverage is correct plus wrong, and conditional accuracy is correct divided by coverage. The Bayes threshold is λ/(1+λ) ; the conditional accuracy of every penalized specialist in the table exceeds its threshold on both datasets.
Dataset (base rate)
λ
Abstained
λ=0 correct
λ=0 wrong
MATH L4–L5 (88.5%)
1
154
33.6%
65.1%
2
200
41.4%
57.7%
5
321
54.8%
44.6%
10
878
75.2%
24.6%
MedMCQA (50.1%)
1
2,020
37.6%
62.3%
2
3,448
43.6%
56.4%
Appendix
Table 11: Performance of the λ=0 model on each specialist’s abstained problems. The λ=0 model’s overall correct rate on the full set is given in parentheses. Every abstained set lies below this base rate, most sharply for the moderate penalties λ=1 and λ=2 .
Unlearning in large language models (LLMs) aims to remove harmful training data while preserving overall utility. However, we find that existing methods often hallucinate, generate abnormal token sequences, or behave inconsistently, raising safety and trust concerns. According to prior literature on LLM honesty, such behaviors are often associated with dishonesty. This motivates us to investigate the notion of honesty in the context of model unlearning. We propose a formal definition of unlearning honesty, which includes: (1) preserving both utility and honesty on retained knowledge, and (2) ensuring effective forgetting while encouraging the model to acknowledge its limitations and respond consistently to questions related to forgotten knowledge. To systematically evaluate the honesty of unlearning, we introduce a suite of metrics that cover utility, honesty on the retained set, effectiveness of forgetting, rejection rate and refusal stability in Q&A and MCQ settings. Evaluating 9 methods across 3 mainstream families shows that all current methods fail to meet these standards. After experimental and theoretical analyses, we present ReVa, a representation-alignment procedure that fine-tunes feature-randomized unlearned models to better acknowledge forgotten knowledge. On Q&A tasks from the forget set, ReVa achieves the highest rejection rate after two rounds of interaction, nearly doubling the performance of the second-best method. Remarkably, It also improves honesty on the retained set. We release our data and code at https://github.com/renjiegu.
Renjie Gu, Jiazhen Du, Yihua Zhang +1
Fudan University · Central South University · Michigan State University
When language models lack relevant knowledge for a given query, they frequently generate plausible responses that can be hallucinations, rather than admitting being agnostic about the answer. Retraining models to reward admitting ignorance can lead to overly conservative behaviors and poor generalization due to scarce evaluation benchmarks. We propose a post hoc framework, Conformal Abstention (CA), adapted from conformal prediction (CP) to determine whether to abstain from answering a query. CA provides finite-sample guarantees on both the probability of participation (i.e., not abstaining) and the probability that the generated response is correct. Importantly, the abstention decision relies on prediction confidence rather than the non-conformity scores used in CP, which are intractable for open-ended generation. To better align prediction confidence with the model's ignorance, we introduce a calibration strategy using representation geometry within the model to measure knowledge involvement in shaping the response. Experiments demonstrate that we improve selective answering significantly with 75 percent conditional correctness.
Rui Xu, Yi Chen, Sihong Xie +1
Information Hub, AI Thrust Hong Kong University of Science and Technology (Guangzhou) Guangzhou, Guangdong, China
A model should refuse two different things: answers it would get wrong, and questions it should not answer at all, such as unanswerable ones or ones resting on a false premise. The usual recipe thresholds a single confidence score, which cannot tell these apart. Across five instruction-tuned models from three families (2B to 14B), we find they are separate axes. Ordinary answer-confidence tracks whether an answer is right but is nearly blind to whether the question is answerable; a linear probe on hidden states does the reverse. The blind spot does not shrink with scale. It is worst on naturally occurring false-premise questions (CREPE). There, answer-confidence, P(IK), P(True), and even asking the model outright whether a premise is false all stay near chance, while a hidden-state probe reaches 0.69 to 0.77 AUROC: the model represents a problem it will not report. This turns out to be fixable. Instructing a model to check premises backfires, because it then disputes sound and false premises alike (57% false challenges), unable to tell them apart; routing the same instruction with the probe roughly triples challenge precision. We turn the two axes into a calibrated policy that answers only when an answerability score and a correctness score each clear a separately certifies behave differently: the unanswerable-answer rate is controllable at every scale, while the wrong-answer rate is capped by model accuracy, so the guarantee tightens as threshold policy certifies both budgets at 0.75 coverage of correct answers, against 0.31 for a single threshold; at 14B it is the only policy that certifies at all.