When Guessing is Rewarded: Rethinking Language Model Evaluation with Distributional Uncertainty Scoring
Organizations: Aleph Alpha Research
Abstract
Standard language model evaluation assigns scores to single predicted answers, rewarding high-confidence responses regardless of how residual probability mass is distributed over alternative options. This creates a systematic pressure toward overconfident guessing: under accuracy-based schemes, a model maximises its expected score by always committing to an answer rather than abstaining, even when its uncertainty is high. While penalty-based approaches partially address this by raising the confidence threshold for strategic guessing, they still treat all sub-threshold responses identically, ignoring a fundamental distinction in how models can express uncertainty - for example between hedging toward incorrect answers versus hedging toward "I don't know" responses. This paper introduces a novel evaluation metric to solve this problem of not considering a model's entire probability distribution over answer choices. The metric naturally distinguishes between harmful overconfidence in wrong answers and uncertainty expressed through abstention, providing scores in an interpretable default range. Through theoretical analysis and illustrative examples, the metric is shown to offer a more nuanced and aligned evaluation paradigm that incentivises models to express genuine uncertainty rather than guessing. Adapting 12 existing evaluation benchmarks to the metric's variants and measuring performance on six language models shows that for half of the tested benchmarks scores are negative across all tested models, indicating significant tendencies towards hallucination.
Figures & tables
| DialoGPT-Medium | Llama3.2 3B Instruct | Llama TFree HAT Pretrained 7B DPO | Mistral 7B Instruct v0.3 | Llama3.1 8B Instruct | DeepSeek R1 0528 Qwen3 8B | |
|---|---|---|---|---|---|---|
| ARC | -24.5 0.3 | -15.1 0.2 | -12.1 0.3 | -10.6 0.3 | -13.9 0.3 | -19.6 0.3 |
| COPA | 0.6 0.4 | 2.9 0.4 | 3.4 0.4 | 10.2 1.0 | 3.0 0.3 | 3.2 1.1 |
| GPQA | -21.6 0.4 | -44.6 0.7 | -43.0 1.1 | -44.7 1.6 | -41.9 1.0 | -39.3 0.5 |
| HellaSwag | -34.0 0.2 | -26.2 0.1 | -23.2 0.2 | -19.0 0.2 | -27.8 0.2 | -37.2 0.3 |
| MMLU | -33.8 0.5 | 0.8 0.5 | 14.0 2.1 | 18.3 2.6 | 19.0 1.9 | 9.3 1.6 |
| MMLU Pro | -61.5 0.3 | -56.8 0.3 | -38.8 1.9 | -34.6 2.6 | -38.9 1.8 | -40.3 1.4 |
| Benchmark | Accuracy | C-weighted acc. | Ternary |
|---|---|---|---|
| ARC | -28.9 | -19.7 | -23.8 |
| COPA | -14.8 | -4.6 | -8.6 |
| GPQA | -66.5 | -49.6 | -7.4 |
| HellaSwag | -78.8 | -41.4 | -49.6 |
| MMLU | -51.9 | -37.8 | -12.2 |
| MMLU Pro | -77.6 | -63.9 | -14.6 |
Appendix figures & tables7 assets
Supplementary material from the paper’s appendix.
Appendix
| Asset | License (summary) |
|---|---|
| transformers library | Apache License 2.0 (repository LICENSE ) |
| eval-framework | Apache License 2.0 (repository LICENSE ) |
| Model checkpoints (Hugging Face model cards) | |
| microsoft/DialoGPT-medium | MIT |
| meta-llama/Llama-3.2-3B-Instruct | Llama 3.2 Community License Agreement |
| meta-llama/Llama-3.1-8B-Instruct | Llama 3.1 Community License Agreement |
| Desired threshold | Loadings ( ) |
|---|---|
| Accuracy | if is , , else |
|---|---|
| C-weighted acc. | if is , , else |
| Ternary score | let ; if , ; if , ; otherwise |
| DialoGPT-Medium | Llama3.2 3B Instruct | Llama TFree HAT Pretrained 7B DPO | Mistral 7B Instruct v0.3 | Llama3.1 8B Instruct | DeepSeek R1 0528 Qwen3 8B | |
|---|---|---|---|---|---|---|
| ARC | 7.3 0.8 | 10.0 0.5 | 4.1 0.6 | 16.6 1.2 | 14.0 1.1 | 25.4 1.4 |
| COPA | 4.0 2.0 | 0.0 0.0 | 0.0 0.0 | 53.0 5.0 | 0.0 0.0 | 55.0 5.0 |
| GPQA | 4.6 0.9 | 32.1 2.0 | 28.6 1.9 | 30.3 2.0 | 32.1 2.0 | 36.5 2.1 |
| HellaSwag | 28.8 1.4 | 48.4 0.5 | 62.8 1.5 | 34.7 1.5 | 69.0 1.5 | 61.8 1.5 |
| MMLU | 18.6 1.2 | 62.1 0.4 | 61.1 1.4 | 59.7 1.5 | 67.8 1.4 | 69.8 1.4 |
| MMLU Pro | 7.8 0.8 | 31.3 0.4 | 37.2 1.5 | 34.3 1.6 | 40.5 1.5 | 43.8 1.5 |
| DialoGPT-Medium | Llama3.2 3B Instruct | Llama TFree HAT Pretrained 7B DPO | Mistral 7B Instruct v0.3 | Llama3.1 8B Instruct | DeepSeek R1 0528 Qwen3 8B | |
|---|---|---|---|---|---|---|
| ARC | 1.7 0.2 | 3.0 0.2 | 1.2 0.2 | 5.5 0.4 | 4.3 0.3 | 6.7 0.4 |
| COPA | 1.5 0.7 | 0.0 0.0 | 0.0 0.0 | 24.8 2.4 | 0.0 0.0 | 24.8 2.3 |
| GPQA | 1.2 0.2 | 10.9 0.7 | 11.6 0.8 | 15.0 1.1 | 12.4 0.8 | 11.7 0.7 |
| HellaSwag | 7.3 0.4 | 11.9 0.1 | 16.4 0.4 | 9.6 0.4 | 17.8 0.4 | 18.2 0.5 |
| MMLU | 6.9 0.4 | 42.0 0.3 | 49.8 1.3 | 55.7 1.4 | 52.6 1.2 | 47.3 1.0 |
| MMLU Pro | 1.4 0.2 | 13.0 0.2 | 23.9 1.1 | 29.0 1.4 | 23.4 1.0 | 21.6 0.9 |
| DialoGPT-Medium | Llama3.2 3B Instruct | Llama TFree HAT Pretrained 7B DPO | Mistral 7B Instruct v0.3 | Llama3.1 8B Instruct | DeepSeek R1 0528 Qwen3 8B | |
|---|---|---|---|---|---|---|
| ARC | -6.9 1.5 | 8.2 0.6 | 3.7 0.7 | 14.0 1.3 | 11.9 1.2 | 15.9 1.8 |
| COPA | 3.0 2.2 | 0.0 0.0 | 0.0 0.0 | 48.0 5.9 | 0.0 0.0 | 24.0 9.0 |
| GPQA | -10.1 1.8 | -35.8 4.0 | -42.8 3.9 | -39.4 3.9 | -35.8 4.0 | -27.0 4.1 |
| HellaSwag | -40.3 2.9 | 27.9 0.8 | 49.5 2.3 | 28.6 1.8 | 40.6 2.8 | 23.6 3.1 |
| MMLU | -40.7 2.3 | 24.3 0.8 | 22.3 2.9 | 19.6 2.9 | 35.6 2.8 | 39.7 2.7 |
| MMLU Pro | -58.2 2.0 | -37.3 0.8 | -25.5 3.0 | -31.4 3.2 | -19.0 3.0 | -11.9 3.1 |
| DialoGPT-Medium | Llama3.2 3B Instruct | Llama TFree HAT Pretrained 7B DPO | Mistral 7B Instruct v0.3 | Llama3.1 8B Instruct | DeepSeek R1 0528 Qwen3 8B | |
|---|---|---|---|---|---|---|
| ARC | -19.8 0.3 | -12.3 0.2 | -9.9 0.2 | -8.7 0.3 | -11.3 0.3 | -15.9 0.3 |
| COPA | 0.5 0.3 | 2.3 0.3 | 2.7 0.3 | 8.3 0.8 | 2.4 0.3 | 2.6 0.9 |
| GPQA | -17.5 0.4 | -36.2 0.6 | -34.9 0.9 | -36.3 1.3 | -34.0 0.8 | -31.9 0.4 |
| HellaSwag | -27.6 0.2 | -21.3 0.1 | -18.8 0.2 | -15.4 0.2 | -22.6 0.2 | -30.2 0.2 |
| MMLU | -27.4 0.4 | 0.7 0.4 | 11.4 1.7 | 14.9 2.1 | 15.4 1.5 | 7.6 1.3 |
| MMLU Pro | -49.9 0.3 | -46.1 0.3 | -31.5 1.5 | -28.1 2.1 | -31.6 1.5 | -32.7 1.1 |