LLM evaluation is commonly performed either by prompting models to produce answers or by scoring candidate outputs with likelihood-based metrics. In multiple-choice QA, however, standard likelihood-based scoring is still conditioned on the question and answer set, and can therefore leverage the same task-conditioned answer-selection interface used in prompting. We study a complementary protocol based on likelihood ranking of declarative statements constructed from the same question--answer pairs. Across 95 decoder-only models, ranging from 0.1B to 104B parameters, and 10 MCQA datasets, we find a systematic divergence between declarative-statement likelihood ranking and prompted answering. Statement-likelihood accuracy remains comparatively stable across scale, whereas prompted answering improves sharply with scale and instruction-tuning. These results suggest that likelihood preferences over controlled declarative alternatives and task-conditioned answer selection probe distinct aspects of model behavior, and should not be treated as interchangeable.
Figures & tables
Figure 1: Overview of the experimental setting and evaluation. The design enables a comparison between prompting-based evaluation and declarative-statement likelihood ranking.
Dataset (HF Name)
# Question
Avg. Options ( ± Std.)
TIGER-Lab/MMLU-Pro
10,632
9.45 ± 1.50
allenai/ai2_arc
7,698
4.00 ± 0.07
allenai/openbookqa
11,836
4.00 ± 0.00
allenai/qasc
8,685
8.00 ± 0.00
allenai/sciq
13,627
4.00 ± 0.00
maveriq/bigbenchhard
4,539
4.73 ± 3.09
Table 1: Statistics of multiple-choice datasets used in our evaluation, same order as the table: Wang et al. (2024c) ; Clark et al. (2018) ; Mihaylov et al. (2018) ; Khot et al. (2020) ; Welbl et al. (2017) ; Suzgun et al. (2023) ; Talmor et al. (2019) ; Lin et al. (2022) ; Science (2025) ; Sakai et al. (2024) .
Figure 2: Accuracies under prompting (APX) and declarative statement likelihood (APS) settings, across model scales and families. APS shows weak scaling and limited gains from instruction-tuning; APX improves sharply with scale, and larger models outperform APS.
Appendix figures & tables9 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 3: Number of examples and errors in declarative statements generation for each dataset.
Figure 4: Distribution of BERTScore F1 for each statement and question–answer pair.
Figure 5: Distribution of Mean BLEU scores per question.
Figure 6: Distribution of answers in each dataset.
Figure 7: Prompting-based (APX) and standard likelihood-based MCQA (Likelihood-MCQA) accuracy across model scales and families. Regardless of the presence of Instruction fine-tuning, likelihood-based MCQA and APX move in the same general direction, with comparable steepness.
Family
# Models
Base
Instruct
Size Range (B)
Qwen2.5
14
7
7
0.5–72
Qwen3
12
6
6
0.6–31
Llama-3
9
5
4
1–71
Gemma-3
10
5
5
0.3–27
Gemma-2
6
3
3
2–27
OLMo-2
8
4
4
1–32
Appendix
Table 2: Overview of the model families evaluated in this work, reporting the number of base and instruction-tuned variants and their parameter scale.
Figure 8: Prompting-based and declarative-statement likelihood accuracies, for instruction-tuned and base models, as a function of model scale, for each dataset.
Figure 9: Average PPL score assigned to the correct answer by each model, divided by dataset.
Figure 10: Average delta between second lowest and lowest PPL score for declarative statements vs model size.
Probing the capabilities of Large Language Models (LLMs) and building robust solutions for Multiple-Choice Question Answering (MCQA) remain central challenges in natural language understanding. Furthermore, the rapid proliferation of LLMs has created the implicit assumption that more sophisticated prompting techniques yield better performance. Several studies claim such gains, but report them under differing models, prompt wordings and answer-extraction rules, so the gains cannot be attributed to the technique alone. We address this gap with ReasonLab, an evaluation framework in which the prompting technique is a first-class experimental variable alongside the model and the dataset, and which retains every generation for inspection. Using ReasonLab we conduct a controlled study of 8 prompting techniques across 10 MCQA datasets, 27 model configurations and 480,927 evaluations at temperature 0. We find that the prompting technique is a minor determinant of accuracy: on configurations without a reasoning budget the reasoning triggers improve on direct prompting by only 3.92 to 4.69 pp and are indistinguishable from one another, and on configurations with reasoning enabled no technique differs by more than 0.51 pp. Self-Generate is the only technique with a consistent effect, a reduction of 2.95 pp. We further investigate three phenomena: (1) the comparison of models on a common set of datasets, where model size does not predict accuracy, (2) the trade-offs across thinking budgets, where enabling reasoning is worth up to 12.74 pp whereas an eightfold budget increase adds only 0.48 to 2.10 pp, and (3) the variation in dataset difficulty, with 60% of benchmarks below 70% accuracy and a 43.9 pp spread from easiest to hardest. These results suggest that, for MCQA, the prompting technique is a minor lever compared with enabling model reasoning, and that substantial headroom remains.
Are large language models (LLMs) bad at capturing human judgment? Two commonly stated limitations are that LLMs fail to capture full distributions of responses, and that their judgments are unstable across wording variations. We demonstrate simple prompting strategies that mitigate these limitations. Across two datasets--a U.S.-representative set of 144 moral scenarios and 38 moral beliefs from the International Social Survey Programme's Family and Changing Gender Roles module covering 32 countries--we show how simple elicitation techniques help improve AI-human alignment. First, prompting models to report standard deviations and response proportions recovers the full range of human responses better than common strategies. Second, ensuring scenarios are clear to human participants--as reflected in human confusion ratings--boosts model alignment, and LLMs can track human confusion ratings. At the same time, we find that LLMs' estimates of their own error are poorly calibrated, though they can predict human variability relatively well. These results suggest that asking better questions to LLMs can yield better answers.
Danica Dillion, Chen Cecilia Liu, Baihui Wang +5
Complexity Science Hub, Vienna · The Ohio State University · University of Cambridge +2
Multiple-choice question answering (MCQA) benchmarks in NLP use number-right scoring (accuracy), but in educational testing, the scoring scheme, the combination of the response mode models follow and the rule for grading responses, is a key design choice that dictates which abilities to reward. We examine how alternatives to number right change what MCQA measures with six education-inspired schemes that assess abilities beyond accuracy: distractor elimination, abstention, confidence calibration, and self-correction. On LLM benchmarks, these schemes: 1) shift rankings of 31 LLMs beyond rephrased number right prompts; 2) better predict the LLMs users prefer in LLM Arena; and 3) reveal distinct model capabilities, like that GPT-5 rarely abstains and readily self-corrects, while weaker open-weight models often abstain and hesitate to eliminate choices. Given the benefits of alternative scoring schemes, we discuss ways to extend them to tasks beyond MCQA.
Nishant Balepur, Paiheng Xu, Wei Ai +3
University of Maryland · New York University · Nanyang Technological University