Also Small Models Can Reasonably Self-Evaluate Their Confidence
Organizations: Technical University of Munich · Siemens AG · LMU Munich · Munich Center for Machine Learning (MCML)
Abstract
This study systematically evaluates self-evaluation-based uncertainty quantification across different language models of varying sizes on question-answering tasks spanning general to specialized knowledge domains. Using various self-evaluation methods where models judge their own predictions, we examine how model scale and domain specificity affect the quality of self-assessed confidence signals. Our results reveal that while accuracy predictably declines with smaller models and more specialized domains, the reliability of self-evaluated confidence remains largely stable across both dimensions. This independence means the most capable model is not necessarily the best at self-assessing prediction reliability. These findings suggest that smaller models can achieve reasonable self-assessed confidence despite lower accuracy, making them viable for resource-constrained deployments.
Figures & tables
| TQA | MMLU | MedMCQA | ||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| OpenAI GPT-4.1 | Ministral-3-3B | OpenAI GPT-4.1 | Ministral-3-3B | OpenAI GPT-4.1 | Ministral-3-3B | |||||||
| Method | Acc. | CalAUC | Acc. | CalAUC | Acc. | CalAUC | Acc. | CalAUC | Acc. | CalAUC | Acc. | CalAUC |
| P(True) (single) | 64.95 | 70.84 | 36.47 | 57.18 | 55.79 | 62.23 | 34.88 | 66.85 | 42.20 | 63.91 | 20.20 | 65.24 |
| Seq likelihood | 71.89 | 47.92 | 37.58 | 57.00 | 62.61 | 56.54 | 33.12 | 60.44 | 44.90 | 63.29 | 19.80 | 62.13 |
| Seq len-norm likelihood | 71.64 | 52.71 | 35.86 | 56.50 | 62.75 | 59.19 | 33.05 | 63.06 | 45.10 | 66.95 | 19.20 | 69.64 |
| Sample and Select | 73.61 | 50.98 | 38.56 | 56.26 | 64.32 | 51.64 | 34.94 | 54.53 | 45.10 | 55.36 | 19.40 | 54.94 |
Appendix figures & tables2 assets
Supplementary material from the paper’s appendix.
Appendix
| OpenAI GPT-4.1 | GPT-OSS-120B | Qwen3.5-9B | Gemma-4-E4B-it | Ministral-3-3B | |||||||
| Dataset | Method | Acc. | CalAUC | Acc. | CalAUC | Acc. | CalAUC | Acc. | CalAUC | Acc. | CalAUC |
| TQA | P(True) (single gen) | 64.95 | 70.84 | 59.98 | 50.87 | 60.71 | 73.09 | 48.47 | 71.15 | 36.47 | 57.18 |
| Seq likelihood | 71.89 | 47.92 | 60.00 | 62.71 | 58.87 | 56.76 | 50.31 | 55.06 | 37.58 | 57.00 | |
| Seq len-norm likelihood | 71.64 | 52.71 | 60.12 | 69.19 | 59.85 | 63.54 | 50.06 | 61.36 | 35.86 | 56.50 | |
| Sample and Select | 73.61 | 50.98 | 64.42 | 58.45 | 61.57 | 52.60 | 50.92 | 50.08 | 38.56 | 56.26 | |
| Sample and Select w/ nota | 73.37 | 56.59 | 63.19 | 58.31 | 61.08 | 59.77 | 50.80 | 57.75 | 37.21 | 61.78 | |
| Model | Parameters | Checkpoint / Endpoint | Access |
|---|---|---|---|
| GPT-4.1 | undisclosed | openai/gpt-4.1 | OpenAI API |
| GPT-OSS-120B | 120B | openai/gpt-oss-120b | vLLM (HF) |
| Qwen3.5-9B | 9B | Qwen/Qwen3.5-9B | vLLM (HF) |
| Gemma-4-E4B-it | 4B (eff.) | google/gemma-4-e4b-it | vLLM (HF) |
| Ministral-3-3B | 3B | mistralai/Ministral-3B-Instruct-2512 | vLLM (HF) |