Probability is Not Enough: Exploring and Counting Divergent Tokens for Reasoning Uncertainty Quantification in LLMs
Organizations: College of Computer Science and Software Engineering, Shenzhen University · Tsinghua University · Behavioral and Spatial AI Lab, Peking University & Tongji University
Abstract
As the chain-of-thought reasoning capabilities of large language models improve, evaluating and calibrating their reasoning confidence is becoming increasingly important for quantifying the uncertainty of their answers. Current methods for estimating the confidence of large language models are generally based on probabilities of selected key tokens, but the underlying mechanism remains unclear. Our pilot study finds that replacing selected token probabilities with coarse substitutes can also improve calibration, motivating us to further explore effective signals of model confidence. We introduce Divergent Token Confidence (DTC), a framework that estimates confidence by counting tokens at which two models strongly disagree during decoding. DTC identifies these divergent tokens using the Jensen-Shannon divergence between next-token distributions evaluated along the same reasoning trajectory. We find that their count is almost negatively associated with answer accuracy, thereby serving as a simple yet effective signal for uncertainty quantification. DTC supports both white-box and black-box evaluation using auxiliary models, without explicit training and affecting the generation process. Experiments across multiple model families and six mathematical benchmarks demonstrate improved calibration over probability-based and verbalized baselines. Under white-box evaluation, the count-only estimator achieves an average expected calibration error of 13.0%, compared with 32.7%-42.4% for standard full-sequence confidence methods. In black-box settings, it also improves calibration over the original verbalized scores. For example, mean expected calibration error falls from 32.1%-40.2% to 13.7%-16.3% on DeepSeek-V3.2. These findings provide new insights for improving reasoning uncertainty quantification in large language models. The code is released at https://github.com/szu-tera/DTC.git.
Figures & tables
| Reasoning Models Acc Entropy Conf. BaseCal UQAC Verb. DTC (Ours) attn mean † prod lin MATH-500 Qwen2.5 - 7B 76.1 43.6 45.3 32.8 43.7 26.8 12.7 44.2 20.0 13.4 Qwen2.5 - 14B 80.0 43.0 44.9 38.9 43.0 26.8 14.2 34.7 16.2 0 9.9 Qwen2.5 - 32B 82.4 43.7 45.4 39.9 43.8 26.1 16.2 41.6 16.9 0 6.6 Qwen3 - 8B 83.9 38.1 41.5 31.8 36.7 37.1 37.8 47.8 0 7.9 15.6 Qwen3 - 14B 86.5 36.8 40.6 29.9 36.2 27.1 13.2 42.5 0 5.5 11.6 Qwen3 - 32B 83.7 36.1 40.2 29.4 — 38.8 38.6 37.0 0 6.1 11.0 Gemma3 - 12B 84.9 41.5 44.0 37.3 34.8 33.9 29.3 24.3 15.2 13.9 Gemma3 - 27B 89.2 42.3 44.5 38.4 36.0 16.0 0 7.0 33.4 15.0 0 7.1 AMC23 Qwen2.5 - 7B 53.6 43.6 45.3 31.7 44.0 24.7 0 8.2 42.5 15.7 0 6.7 Qwen2.5 - 14B 61.2 42.7 44.7 38.5 43.1 25.3 19.7 35.1 0 9.7 0 9.5 Qwen2.5 - 32B 66.4 43.3 45.1 39.5 44.0 24.9 22.7 38.3 0 9.8 10.8 Qwen3 - 8B 68.6 36.5 40.5 29.4 36.8 49.4 38.5 48.7 0 6.8 0 9.3 Qwen3 - 14B 74.1 34.6 39.2 26.9 35.9 27.2 10.5 35.3 12.3 11.8 Qwen3 - 32B 67.5 33.7 38.5 25.4 — 44.0 37.7 32.3 15.2 0 7.5 Gemma3 - 12B 66.8 40.8 43.6 36.4 35.3 34.5 24.0 25.9 12.7 10.9 Gemma3 - 27B 76.9 41.8 44.3 37.9 36.9 19.9 12.3 26.4 12.4 12.0 AIME24 Qwen2.5 - 7B 12.6 42.8 44.7 30.3 43.3 29.5 19.0 39.5 14.7 11.5 Qwen2.5 - 14B 13.8 41.8 44.0 37.5 42.3 25.1 15.6 31.4 13.1 16.9 Qwen2.5 - 32B 16.9 42.4 44.4 37.9 43.5 23.9 16.7 34.1 10.6 20.4 Qwen3 - 8B 27.7 34.9 39.3 26.5 36.4 40.3 42.9 45.8 14.2 14.9 Qwen3 - 14B 27.1 33.2 38.1 24.4 35.7 31.4 16.2 44.1 20.6 18.1 Qwen3 - 32B 28.3 32.0 37.4 23.7 — 44.2 37.6 25.7 21.6 16.4 Gemma3 - 12B 23.9 39.8 42.8 34.7 34.6 33.2 31.8 23.0 0 9.7 12.6 Gemma3 - 27B 29.0 40.7 43.5 36.2 36.1 29.6 29.4 21.1 0 8.1 24.3 AIME25 Qwen2.5 - 7B 0 9.1 42.4 44.4 31.0 43.1 26.6 16.2 37.3 11.2 20.5 Qwen2.5 - 14B 14.7 42.4 44.4 37.8 42.9 27.5 15.7 31.6 10.8 25.4 Qwen2.5 - 32B 12.2 42.7 44.7 38.3 43.8 30.6 18.4 33.0 0 8.7 28.3 Qwen3 - 8B 22.5 34.2 38.8 25.9 36.0 37.6 43.4 46.9 12.3 0 9.5 Qwen3 - 14B 27.5 32.6 37.8 24.0 35.3 30.8 19.7 53.0 17.1 0 5.3 Qwen3 - 32B 23.7 31.7 37.2 22.5 — 36.3 37.8 25.6 19.1 0 6.9 Gemma3 - 12B 18.8 40.3 43.2 35.5 34.5 32.1 29.4 15.9 16.7 0 7.8 Gemma3 - 27B 25.3 40.6 43.4 35.9 35.4 25.1 22.7 23.1 15.6 11.1 Average 48.0 39.3 42.4 32.7 39.0 30.8 23.6 35.0 13.2 13.0 | PRM (ref.) 0 9.8 0 8.8 0 8.4 0 7.4 0 6.2 0 6.3 0 7.2 0 8.4 0 6.0 0 6.0 0 8.9 11.4 12.0 11.6 10.1 13.1 22.7 23.0 23.4 29.2 29.7 28.8 29.2 30.0 25.2 24.4 27.7 29.6 30.7 30.8 27.2 26.7 18.1 |
| Method | AIME24 | AIME25 | HMMT25 | HMMT26 | ||||||||
| Acc | ECE | AUC | Acc | ECE | AUC | Acc | ECE | AUC | Acc | ECE | AUC | |
| Default CoT | 72.2 | – | – | 60.9 | – | – | 46.1 | – | – | 50.9 | – | – |
| + PRM | – | 39.0 | 88.2 | – | 39.5 | 80.4 | – | 41.1 | 60.4 | – | 38.1 | 75.1 |
| + | – | 11.2 | 74.6 | – | 0 9.4 | 78.7 | – | 12.7 | 67.9 | – | 12.5 | 72.7 |
| + | – | 24.6 | 79.4 | – | 23.9 | 82.1 | – | 29.1 | 75.0 | – | 26.4 | 76.5 |
| Verb. Conf. | 67.4 | 40.8 | 85.5 | 55.3 | 39.6 | 86.8 | 38.3 | 40.3 | 78.2 | 45.0 | 40.1 | 84.3 |
| Method | AIME24 | AIME25 | HMMT25 | HMMT26 | ||||
| ECE | AUC | ECE | AUC | ECE | AUC | ECE | AUC | |
| Verb. Conf. | 40.8 | 85.5 | 39.6 | 86.8 | 40.3 | 78.2 | 40.1 | 84.3 |
| w/ DTC | 0 6.3 | 85.8 | 0 3.7 | 87.7 | 10.5 | 77.8 | 0 6.2 | 84.2 |
| Verb. TopK | 32.0 | 88.4 | 32.6 | 85.7 | 31.6 | 79.3 | 32.0 | 83.6 |
| w/ DTC | 17.0 | 90.3 | 14.6 | 87.3 | 18.8 | 79.1 | 15.7 | 83.6 |
| Verb. PD | 32.6 | 86.2 | 33.2 | 84.9 | 31.5 | 74.3 | 32.3 | 83.1 |
Appendix figures & tables17 assets
Supplementary material from the paper’s appendix.
Appendix
| White-box | Black-box | |
| Benchmarks | MATH-500 ( Lightman et al., 2024 ) AMC23 ( Balunovic et al., 2025 ) AIME24 ( Balunovic et al., 2025 ) AIME25 ( Balunovic et al., 2025 ) | AIME24 ( Balunovic et al., 2025 ) AIME25 ( Balunovic et al., 2025 ) HMMT25 ( Balunovic et al., 2025 ) HMMT26 ( Balunovic et al., 2025 ) |
| Generators | Qwen2.5-7B-Instruct ( Qwen Team, 2024 ) Qwen2.5-14B-Instruct ( Qwen Team, 2024 ) Qwen2.5-32B-Instruct ( Qwen Team, 2024 ) Qwen3-8B ( Yang et al., 2025 ) Qwen3-14B ( Yang et al., 2025 ) Qwen3-32B ( Yang et al., 2025 ) Gemma3-12B-IT ( Gemma Team, 2025 ) Gemma3-27B-IT ( Gemma Team, 2025 ) | DeepSeek-V3.2 ( DeepSeek-AI, 2025 ) Qwen3-30B-A3B-Instruct-2507 ( Yang et al., 2025 ) Qwen3-4B-Instruct-2507 ( Yang et al., 2025 ) |
| Auxiliaries | Qwen2.5-1.5B-Instruct ( Qwen Team, 2024 ) Qwen3-1.7B ( Yang et al., 2025 ) Gemma3-4B-IT ( Gemma Team, 2025 ) | Qwen2.5-7B-Instruct ( Qwen Team, 2024 ) Qwen2.5-1.5B-Instruct ( Qwen Team, 2024 ) |
| Sample number per question | on MATH-500 on AMC23 on AIME24 on AIME25 | for DeepSeek-V3.2 for Qwen3-30B-A3B-Instruct-2507 for Qwen3-4B-Instruct-2507 |
| Reasoning Models Acc Entropy Conf. BaseCal UQAC Verb. DTC (Ours) attn mean † prod lin MATH-500 Qwen2.5 - 7B 76.1 67.0 66.5 61.4 58.5 59.2 62.0 74.7 80.4 83.8 Qwen2.5 - 14B 80.0 65.7 65.0 65.0 56.0 56.6 64.9 74.8 81.9 85.0 Qwen2.5 - 32B 82.4 66.7 65.8 65.5 50.6 59.7 65.0 83.8 82.7 85.8 Qwen3 - 8B 83.9 75.0 73.4 72.4 52.3 59.1 61.6 57.9 79.3 73.9 Qwen3 - 14B 86.5 74.3 72.8 72.0 53.2 66.1 75.6 61.8 79.4 75.6 Qwen3 - 32B 83.7 73.6 71.6 70.9 — 58.8 63.1 63.3 80.6 78.9 Gemma3 - 12B 84.9 61.9 60.3 62.7 40.9 61.9 58.1 84.1 81.9 84.0 Gemma3 - 27B 89.2 63.6 61.8 63.3 38.9 76.5 81.7 87.2 85.5 87.9 AMC23 Qwen2.5 - 7B 53.6 72.4 71.8 67.9 66.6 63.9 68.1 67.1 81.3 81.2 Qwen2.5 - 14B 61.2 66.9 66.4 67.7 62.3 57.1 68.5 73.0 79.2 80.3 Qwen2.5 - 32B 66.4 67.1 66.3 69.8 60.8 62.2 68.6 73.7 77.6 78.0 Qwen3 - 8B 68.6 73.8 72.0 71.6 56.2 41.1 47.7 57.8 78.7 76.0 Qwen3 - 14B 74.1 73.5 71.5 72.6 56.6 66.1 72.9 68.8 79.2 75.3 Qwen3 - 32B 67.5 75.3 73.4 74.4 — 55.1 57.2 67.7 80.3 78.6 Gemma3 - 12B 66.8 69.7 68.6 70.0 54.2 61.2 61.9 78.7 79.9 78.7 Gemma3 - 27B 76.9 68.5 67.1 68.5 47.8 72.7 76.4 85.6 82.6 80.4 AIME24 Qwen2.5 - 7B 12.6 79.0 78.0 74.4 75.2 55.1 61.5 79.0 84.4 75.6 Qwen2.5 - 14B 13.8 71.3 71.0 75.2 65.5 61.4 70.1 77.8 76.5 74.2 Qwen2.5 - 32B 16.9 75.6 75.2 79.3 72.2 67.5 73.5 84.1 80.0 70.4 Qwen3 - 8B 27.7 71.3 69.9 72.1 58.4 53.5 56.4 65.5 74.1 65.7 Qwen3 - 14B 27.1 70.8 69.4 70.6 61.4 64.9 72.6 63.9 74.3 67.0 Qwen3 - 32B 28.3 71.2 69.9 73.4 — 52.1 52.5 74.7 76.2 69.6 Gemma3 - 12B 23.9 81.0 80.1 79.8 71.1 64.8 60.9 74.8 81.8 66.4 Gemma3 - 27B 29.0 78.5 78.1 78.3 68.2 66.4 65.1 87.8 80.2 64.0 AIME25 Qwen2.5 - 7B 0 9.1 79.3 78.9 77.8 78.7 67.7 71.1 69.6 76.5 64.3 Qwen2.5 - 14B 14.7 75.1 74.7 75.9 71.8 68.5 75.5 74.4 75.9 62.2 Qwen2.5 - 32B 12.2 79.0 79.1 79.4 75.4 52.4 77.5 72.1 76.7 60.1 Qwen3 - 8B 22.5 78.9 77.4 77.2 67.9 55.1 56.2 63.3 82.4 70.9 Qwen3 - 14B 27.5 82.5 81.3 82.8 73.8 67.7 84.1 52.1 89.1 81.8 Qwen3 - 32B 23.7 80.3 78.8 81.3 — 61.2 61.7 74.8 86.8 79.9 Gemma3 - 12B 18.8 84.9 84.1 84.4 69.0 64.8 66.7 85.7 88.9 76.3 Gemma3 - 27B 25.3 86.9 85.8 86.4 67.2 70.9 74.8 91.3 89.8 79.0 Average 48.0 73.8 72.7 73.2 61.8 61.6 66.7 73.5 80.8 75.3 | PRM (ref.) 95.4 94.3 94.4 90.9 89.8 91.3 93.6 93.2 90.8 90.5 87.2 90.6 89.5 89.0 86.6 91.1 96.2 89.7 89.6 89.0 88.6 91.7 93.8 90.8 88.6 88.3 83.8 87.5 87.5 88.1 89.6 91.1 90.4 |
| Method | AIME24 | AIME25 | HMMT25 | HMMT26 | ||||||||
| Acc | ECE | AUC | Acc | ECE | AUC | Acc | ECE | AUC | Acc | ECE | AUC | |
| Default CoT | 61.4 | – | – | 45.6 | – | – | 30.3 | – | – | 34.7 | – | – |
| + PRM | – | 36.0 | 82.9 | – | 36.6 | 80.3 | – | 37.8 | 53.3 | – | 33.4 | 78.0 |
| + | – | 11.3 | 72.9 | – | 0 9.0 | 81.5 | – | 13.1 | 69.6 | – | 13.0 | 74.6 |
| + | – | 16.3 | 75.8 | – | 10.7 | 84.8 | – | 15.4 | 78.0 | – | 14.3 | 79.0 |
| Verb. Conf. | 61.8 | 45.6 | 80.8 | 46.6 | 45.9 | 85.6 | 27.3 | 45.6 | 81.9 | 33.0 | 45.1 | 79.8 |
| Method | AIME24 | AIME25 | HMMT25 | HMMT26 | ||||||||
| Acc | ECE | AUC | Acc | ECE | AUC | Acc | ECE | AUC | Acc | ECE | AUC | |
| Default CoT | 73.4 | – | – | 59.8 | – | – | 42.4 | – | – | 43.5 | – | – |
| + PRM | – | 37.5 | 86.0 | – | 37.8 | 78.4 | – | 38.7 | 53.6 | – | 34.6 | 77.4 |
| + | – | 15.9 | 70.2 | – | 12.3 | 75.8 | – | 13.1 | 72.1 | – | 0 9.6 | 77.8 |
| + | – | 18.2 | 75.1 | – | 16.6 | 79.6 | – | 19.5 | 78.7 | – | 17.1 | 80.5 |
| Verb. Conf. | 75.7 | 44.5 | 83.9 | 60.6 | 44.3 | 80.8 | 41.8 | 42.9 | 85.3 | 43.9 | 43.9 | 79.2 |
| Method | AIME24 | AIME25 | HMMT25 | HMMT26 | ||||
| ECE | AUC | ECE | AUC | ECE | AUC | ECE | AUC | |
| Verb. Conf. | 44.5 | 83.9 | 44.3 | 80.8 | 42.9 | 85.3 | 43.9 | 79.2 |
| w/ DTC | 17.8 | 87.0 | 15.7 | 83.8 | 13.0 | 87.9 | 17.4 | 85.9 |
| Verb. TopK | 40.2 | 88.2 | 40.6 | 88.7 | 41.2 | 85.4 | 41.6 | 83.0 |
| w/ DTC | 11.1 | 83.0 | 10.4 | 84.9 | 0 5.4 | 86.5 | 0 6.1 | 86.4 |
| Verb. PD | 33.5 | 75.5 | 34.1 | 69.6 | 31.1 | 77.3 | 34.9 | 72.3 |
| Method | AIME24 | AIME25 | HMMT25 | HMMT26 | ||||
| ECE | AUC | ECE | AUC | ECE | AUC | ECE | AUC | |
| Verb. Conf. | 45.6 | 80.8 | 45.9 | 85.6 | 45.6 | 81.9 | 45.1 | 79.8 |
| w/ DTC | 30.8 | 82.0 | 29.4 | 87.0 | 31.0 | 81.6 | 31.1 | 81.2 |
| Verb. TopK | 47.0 | 68.3 | 47.0 | 75.3 | 46.8 | 74.0 | 45.5 | 77.3 |
| w/ DTC | 33.7 | 68.6 | 32.8 | 75.5 | 34.3 | 73.8 | 30.9 | 76.9 |
| Verb. PD | 46.9 | 69.9 | 48.1 | 73.7 | 43.9 | 75.5 | 42.9 | 70.9 |