Likelihood Ranking doesn't Scale Like Prompting in LLMs
Organizations: CoLingLab, Department of Philology, Literature and Linguistics, University of Pisa · Department of Computer Science, University of Pisa
Abstract
LLM evaluation is commonly performed either by prompting models to produce answers or by scoring candidate outputs with likelihood-based metrics. In multiple-choice QA, however, standard likelihood-based scoring is still conditioned on the question and answer set, and can therefore leverage the same task-conditioned answer-selection interface used in prompting. We study a complementary protocol based on likelihood ranking of declarative statements constructed from the same question--answer pairs. Across 95 decoder-only models, ranging from 0.1B to 104B parameters, and 10 MCQA datasets, we find a systematic divergence between declarative-statement likelihood ranking and prompted answering. Statement-likelihood accuracy remains comparatively stable across scale, whereas prompted answering improves sharply with scale and instruction-tuning. These results suggest that likelihood preferences over controlled declarative alternatives and task-conditioned answer selection probe distinct aspects of model behavior, and should not be treated as interchangeable.
Figures & tables
| Dataset (HF Name) | # Question | Avg. Options ( Std.) |
|---|---|---|
| TIGER-Lab/MMLU-Pro | 10,632 | 9.45 1.50 |
| allenai/ai2_arc | 7,698 | 4.00 0.07 |
| allenai/openbookqa | 11,836 | 4.00 0.00 |
| allenai/qasc | 8,685 | 8.00 0.00 |
| allenai/sciq | 13,627 | 4.00 0.00 |
| maveriq/bigbenchhard | 4,539 | 4.73 3.09 |
Appendix figures & tables9 assets
Supplementary material from the paper’s appendix.
Appendix
| Family | # Models | Base | Instruct | Size Range (B) |
|---|---|---|---|---|
| Qwen2.5 | 14 | 7 | 7 | 0.5–72 |
| Qwen3 | 12 | 6 | 6 | 0.6–31 |
| Llama-3 | 9 | 5 | 4 | 1–71 |
| Gemma-3 | 10 | 5 | 5 | 0.3–27 |
| Gemma-2 | 6 | 3 | 3 | 2–27 |
| OLMo-2 | 8 | 4 | 4 | 1–32 |