cs.CLOct 5, 2025

When Guessing is Rewarded: Rethinking Language Model Evaluation with Distributional Uncertainty Scoring

Authors: Thomas F Burns

Organizations: Aleph Alpha Research

Abstract

Standard language model evaluation assigns scores to single predicted answers, rewarding high-confidence responses regardless of how residual probability mass is distributed over alternative options. This creates a systematic pressure toward overconfident guessing: under accuracy-based schemes, a model maximises its expected score by always committing to an answer rather than abstaining, even when its uncertainty is high. While penalty-based approaches partially address this by raising the confidence threshold for strategic guessing, they still treat all sub-threshold responses identically, ignoring a fundamental distinction in how models can express uncertainty - for example between hedging toward incorrect answers versus hedging toward "I don't know" responses. This paper introduces a novel evaluation metric to solve this problem of not considering a model's entire probability distribution over answer choices. The metric naturally distinguishes between harmful overconfidence in wrong answers and uncertainty expressed through abstention, providing scores in an interpretable default range. Through theoretical analysis and illustrative examples, the metric is shown to offer a more nuanced and aligned evaluation paradigm that incentivises models to express genuine uncertainty rather than guessing. Adapting 12 existing evaluation benchmarks to the metric's variants and measuring performance on six language models shows that for half of the tested benchmarks scores are negative across all tested models, indicating significant tendencies towards hallucination.

Figures & tables

Appendix figures & tables7 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Calibration Is Not Enough: Evaluating Confidence Estimation Under Language Variations

    Jan 12, 2026Yuxi Xia, Dennis Ulmer, Terra Blevins +3Confidence Estimation

  2. Position: Evaluation Scores Are Perishable Knowledge Claims

    Jul 28, 2026Sankalp Gilda, Shlok GildaConfidence CalibrationPosition

  3. Geometric Uncertainty for Detecting and Correcting Hallucinations in LLMs

    Sep 17, 2025Edward Phillips, Sean Wu, Soheila Molaei +3Large Language Model UncertaintyLarge Language Model Hallucination