Honesty over Accuracy: Trustworthy Language Models through Reinforced Hesitation
Organizations: Toyota Technological Institute at Chicago · University of California, San Diego
Abstract
Modern language models fail a fundamental requirement of trustworthy intelligence: knowing when not to answer. Despite achieving impressive accuracy on benchmarks, these models produce confident hallucinations, even when wrong answers carry catastrophic consequences. Our evaluations on GSM8K, MedQA and GPQA show frontier models almost never abstain despite explicit warnings of severe penalties, suggesting that prompts cannot override training that rewards any answer over no answer. As a remedy, we propose Reinforced Hesitation (RH): a modification to Reinforcement Learning from Verifiable Rewards (RLVR) to use ternary rewards (+1 correct, 0 abstention, - error) instead of binary. Controlled experiments on logic puzzles reveal that varying produces distinct models along a Pareto frontier, where each training penalty yields the optimal model for its corresponding risk regime: low penalties produce aggressive answerers, high penalties conservative abstainers. The same frontier holds on MATH Levels 4--5 and on medical QA, where it transfers to an unseen dataset. We then introduce two inference strategies that exploit trained abstention as a coordination signal: cascading routes queries through models with decreasing risk tolerance, while self-cascading re-queries the same model on abstention. Both outperform majority voting with lower computational cost. These results establish abstention as a first-class training objective that transforms ``I don't know'' from failure into a coordination signal, enabling models to earn trust through calibrated honesty about their limits.
Figures & tables
Appendix figures & tables27 assets
Supplementary material from the paper’s appendix.
Appendix
| Method | Uncertainty signal | Training objective | Relation to RH |
|---|---|---|---|
| RH (ours) | Explicit answer/IDK action | Ternary utility | Learns fixed-risk specialists and uses abstention as an inference-routing signal. |
| TruthRL ( Wei et al., 2026 ) | Explicit abstention | Fixed ternary abstention reward | Closest to the member of RH in verifiable-answer settings. |
| RLCR ( Damani et al., 2026 ) | Verbalized confidence | Correctness plus Brier-style calibration | Direct confidence-threshold baseline; we threshold at . |
| BCRL ( Wu et al., 2026 ) | Verbalized confidence, thresholded post hoc at a user-specified risk level | Proper scoring rules; also a prompted- ternary reward | Adjacent single-model approach to user-specified risk; its prompted- variant did not adapt to . |
| Rewarding Doubt ( Bani-Harouni et al., 2026 ) | Verbalized integer confidence (0 to 10) | Logarithmic scoring rule on confidence, with answers held fixed | Complementary direction focused on calibrated confidence expression. |
| Suite | Model and data | Reward variants |
|---|---|---|
| Knights & Knaves | Qwen/Qwen3-1.7B ( Qwen Team, 2025 ) ; 80,000 training puzzles and 10,000 test puzzles. | Fixed RH penalties . |
| MATH | Qwen3.5-4B; MATH Levels 4–5 with 3,994 training problems and 2,538 validation problems. | Fixed RH penalties , prompt-conditioned dynamic penalty over the same set, expected-penalty ablation at , and an RLCR baseline ( Damani et al., 2026 ) . |
| Setting | Value |
|---|---|
| Shared settings | |
| Training pipeline | verl ( Sheng et al., 2025 ) with SGLang rollouts ( Zheng et al., 2024 ) ; MATH uses the same verl/DrGRPO stack. |
| Rollout sampling | Temperature 1.0 with nucleus sampling (top-p=1.0), 8 rollouts per prompt. |
| Validation decoding | Deterministic decoding, temperature 0, . |
| Response limit | 4096 tokens. |
| Knights & Knaves | |
| Correct | IDK | Wrong | Format err. | Cond. acc. | |
|---|---|---|---|---|---|
| 0 | 88.48% | 0.08% | 11.44% | 1.59% | 88.6% |
| 1 | 87.61% | 6.13% | 6.25% | 0.21% | 93.3% |
| 2 | 85.78% | 8.16% | 6.05% | 0.11% | 93.4% |
| 5 | 81.81% | 12.96% | 5.23% | 0.00% | 94.0% |
| 10 | 62.85% | 34.39% | 2.75% | 0.04% | 95.8% |
| Deployed model | Eval | Eval | Eval | Eval | Eval |
|---|---|---|---|---|---|
| Specialist | 0.890 | 0.781 | 0.671 | 0.344 | -0.202 |
| Specialist | 0.883 | 0.818 | 0.753 | 0.558 | 0.233 |
| Specialist | 0.868 | 0.809 | 0.749 | 0.571 | 0.273 |
| Specialist | 0.841 | 0.785 | 0.729 | 0.561 | 0.282 |
| Specialist | 0.640 | 0.615 | 0.591 | 0.516 | 0.392 |
| RLCR + threshold | Matching RH specialist | |||||||
|---|---|---|---|---|---|---|---|---|
| Correct | Wrong | IDK | Correct | Wrong | IDK | |||
| 0.00 | 88.7 | 11.3 | 0.0 | 88.5 | 11.4 | 0.1 | -0.3 | |
| 0.50 | 87.3 | 7.4 | 5.3 | 87.6 | 6.3 | 6.1 | +1.5 | |
| 0.67 | 80.9 | 5.8 | 13.2 | 85.8 | 6.1 | 8.2 | +4.4 | |
| 0.83 | 80.9 | 5.8 | 13.3 | 81.8 | 5.2 | 13.0 | +4.0 | |
| 0.91 | 42.7 | 2.6 | 54.7 | 62.9 | 2.8 | 34.4 | +18.6 | |
| Cascade order | Correct | IDK | Wrong | Format err. | Avg. queries |
|---|---|---|---|---|---|
| 62.85 | 34.39 | 2.75 | 0.04 | 1.00 | |
| 82.62 | 12.01 | 5.37 | 0.05 | 1.34 | |
| 86.69 | 7.08 | 6.23 | 0.07 | 1.46 | |
| 88.54 | 4.60 | 6.86 | 0.14 | 1.53 | |
| 89.97 | 0.07 | 9.96 | 0.86 | 1.58 |
| Method | Budget | Correct | IDK | Wrong |
|---|---|---|---|---|
| Single | 1 | 88.48 | 0.08 | 11.44 |
| Majority vote | 4 | 90.20 | 0.07 | 9.73 |
| Majority vote | 8 | 90.80 | 0.06 | 9.13 |
| Majority vote | 16 | 91.05 | 0.08 | 8.87 |
| Full cascade | 1.58 avg. | 89.97 | 0.07 | 9.96 |
| Example | Correct at =0 | Correct at =10 | IDK at =10 |
|---|---|---|---|
| 1 | 100.0% | 28.1% | 71.9% |
| 2 | 100.0% | 31.2% | 68.8% |
| 3 | 98.4% | 21.9% | 78.1% |
| 4 | 87.5% | 0.0% | 100.0% |
| 5 | 85.9% | 0.0% | 100.0% |
| MedMCQA (in-domain) | MedQA (transfer) | ||||||
|---|---|---|---|---|---|---|---|
| Threshold | Coverage | Wrong | Cond. acc. | Coverage | Wrong | Cond. acc. | |
| 0 | n/a | 99.98% | 49.91% | 50.1% | 100.0% | 47.50% | 52.5% |
| 1 | 50.0% | 53.76% | 21.87% | 59.3% | 75.60% | 32.78% | 56.6% |
| 2 | 66.7% | 21.06% | 5.17% | 75.5% | 26.02% | 7.08% | 72.8% |
| 5 | 83.3% | 7.25% | 0.94% | 87.0% | 5.61% | 0.80% | 85.7% |
| Dataset (base rate) | Abstained | correct | wrong | |
| MATH L4–L5 (88.5%) | 1 | 154 | 33.6% | 65.1% |
| 2 | 200 | 41.4% | 57.7% | |
| 5 | 321 | 54.8% | 44.6% | |
| 10 | 878 | 75.2% | 24.6% | |
| MedMCQA (50.1%) | 1 | 2,020 | 37.6% | 62.3% |
| 2 | 3,448 | 43.6% | 56.4% |