Multiple-choice question answering (MCQA) is commonly used to evaluate large language models under the assumption that one of the provided options is correct, typically using answer-selection accuracy. However, in real deployments, users or retrieval systems may provide invalid option sets in which none of the listed choices is correct, and selecting one of them may incur downstream cost. We study this setting as penalty-framed no-valid-option MCQA. Using the mathematics subset of MMLU-Pro, we remove the labeled correct option, allow models to either choose a remaining option or output ABSTAIN, and penalize invalid forced-choice responses. We further introduce correct-conditioned analysis, evaluating abstention only on instances that the model originally answered correctly. Experiments show that high MCQA accuracy does not fully guarantee abstention reliability: even under explicit no-valid-option-aware instructions and penalty-based scoring, models still produce invalid forced-choice responses for a subset of originally correct instances. These results show that penalty-framed no-valid-option MCQA reveals an aspect of model reliability not captured by standard answer-selection accuracy.
Figures & tables
Figure 1: Original MCQA accuracy and no-valid-option abstention rate. AR is reported under the main setting: CoT prompting, no-valid-option-aware instruction, invalid forced-choice penalty −1 , robust parsing. Values are 3-run means on the filtered MMLU-Pro Mathematics subset (1,118 instances).
Dataset
Model
Acc ↑
AR ↑
IFR ↓
CAR ↑
CIFR ↓
MMLU-Pro Math
GPT-5-mini
97.5
85.6
14.4
87.2
12.8
GPT-5-nano
94.1
81.1
18.9
84.5
15.5
GPT-4.1-mini
95.3
69.0
31.0
72.4
27.6
GPT-4.1-nano
78.1
47.8
52.2
54.5
45.5
Gemini-3.0-Flash
96.2
90.1
9.9
90.9
9.1
Gemini-3.1-Flash-Lite
96.5
83.8
16.2
86.4
13.6
Table 1: Main post-exclusion results under CoT prompting, robust parsing, and the no-valid-option-aware instruction with invalid forced-choice penalty −1 . AR/IFR are computed over all no-valid-option instances, while CAR/CIFR are computed on each model’s stable-correct subset for the corresponding dataset. Subset sizes are reported in Appendix F , Table 6 . All values are 3-run means reported as percentages; standard deviations are reported in Appendix C , Table 5 .
Scoring-rule variation
Abstention-instruction variation
Model
−0.5
−1
−2
Penalty-only
Uncertainty-aware
No-valid-option-aware
GPT-5-mini
13.1 ±0.4
12.8 ±0.7
13.2 ±0.8
23.5 ±0.3
21.5 ±1.1
12.8 ±0.7
GPT-5-nano
15.2 ±0.4
15.5 ±1.0
14.8 ±0.6
23.4 ±0.8
20.8 ±1.2
15.5 ±1.0
GPT-4.1-mini
27.8 ±0.5
27.6 ±0.6
26.6 ±0.2
41.1 ±0.8
48.2 ±0.1
27.6 ±0.6
GPT-4.1-nano
46.8 ±0.7
45.5 ±0.6
46.4 ±0.5
75.7 ±0.6
65.0 ±0.4
45.5 ±0.6
Gemini-3.0-Flash
10.2 ±0.3
9.1 ±0.6
9.9 ±1.1
31.2 ±0.5
32.5 ±1.4
9.1 ±0.6
Table 2: Stable-correct CIFR on the filtered MMLU-Pro Mathematics subset ( N=1,118 ) under scoring-rule and instruction ablations. In the scoring-rule block, the no-valid-option-aware instruction is fixed and only the invalid forced-choice penalty varies. In the instruction block, +4 / −1 / 0 scoring is fixed and only the abstention instruction varies. CIFR values are reported as percentages in mean ± standard deviation format over three independent runs. Lower CIFR is better.
Appendix figures & tables14 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 2: Construction of penalty-framed no-valid-option MCQA. The labeled correct option is removed from an original MCQA instance. The desired behavior is to output Abstain ; selecting any remaining option is treated as an invalid forced-choice response and incurs a penalty.
Model name in paper
API model identifier
Provider
Runtime setting
GPT-5-mini
gpt-5-mini-2025-08-07
OpenAI
reasoning_effort=low
GPT-5-nano
gpt-5-nano-2025-08-07
OpenAI
reasoning_effort=low
GPT-4.1-mini
gpt-4.1-mini-2025-04-14
OpenAI
temperature=0.0
GPT-4.1-nano
gpt-4.1-nano-2025-04-14
OpenAI
temperature=0.0
Gemini-3.0-Flash
gemini-3-flash-preview
Google Vertex AI
thinking_level=low
Gemini-3.1-Flash-Lite
gemini-3.1-flash-lite
Google Vertex AI
thinking_level=low
Appendix
Table 3: Model identifiers and runtime settings used in our experiments. The six core models are used in the prompting, scoring-rule, and parser-sensitivity analyses. GPT-5, Claude Sonnet 5, and Qwen3-8B are additionally evaluated under the main setting. Gemini 2.5-family models are reported only as supplementary diagnostics because of weaker completion and parsing reliability under the current evaluation pipeline.
Exclusion reason
MMLU-Pro Math
Multi-domain MMLU
Label error
47
26
Alternative correct option
72
34
Total excluded
119
60
Appendix
Table 4: Exclusion counts by reason. Alternative correct options are options that could remain valid after removal of the labeled correct option.
Dataset
Model
Set
N
Acc ↑
AR ↑
IFR ↓
CAR ↑
CIFR ↓
GPT-5-mini
Unfiltered
1237
92.3 ±0.6
79.8 ±0.6
20.2 ±0.6
84.3 ±0.6
15.7 ±0.6
Filtered
1118
97.5 ±0.5
85.6 ±0.7
14.4 ±0.7
87.2 ±0.7
12.8 ±0.7
GPT-5-nano
Unfiltered
1237
89.6 ±0.4
76.4 ±0.8
23.6 ±0.8
81.9 ±0.9
18.1 ±0.9
Filtered
1118
94.1 ±0.4
81.1 ±1.0
18.9 ±1.0
84.5 ±1.0
15.5 ±1.0
GPT-4.1-mini
Unfiltered
1237
91.3 ±0.2
64.3 ±0.6
35.7 ±0.6
69.3 ±0.5
30.7 ±0.5
Filtered
1118
95.3 ±0.3
69.0 ±0.6
31.0 ±0.6
72.4 ±0.6
27.6 ±0.6
Appendix
Table 5: Unfiltered and filtered results on MMLU-Pro Mathematics and multi-domain MMLU under CoT prompting, the no-valid-option-aware instruction, invalid forced-choice penalty −1 , and robust parsing. N is the number of dataset instances; Acc is computed over completed original MCQA responses in each run. AR/IFR are computed over evaluation-eligible no-valid-option responses, and CAR/CIFR over evaluation-eligible responses within each model’s stable-correct subset. Values are percentages reported as mean ± standard deviation over three independent runs.
Dataset
Model
Stable-correct count
Stable-correct (%)
MMLU-Pro Math
GPT-5-mini
1068
95.5
GPT-5-nano
1004
89.8
GPT-4.1-mini
1032
92.3
GPT-4.1-nano
726
64.9
Gemini-3.0-Flash
1025
91.7
Gemini-3.1-Flash-Lite
1046
93.6
Appendix
Table 6: Model-specific stable-correct subset sizes under CoT prompting. An instance is included if the model answers the original MCQA instance correctly in all three runs. Percentages use the corresponding filtered dataset size: 1,118 for MMLU-Pro Mathematics and 340 for multi-domain MMLU.
Model
Acc ↑
AR ↑
IFR ↓
CAR ↑
CIFR ↓
GPT-5-mini
97.3 ±0.1
84.5 ±1.3
15.5 ±1.3
86.6 ±1.1
13.4 ±1.1
GPT-5-nano
93.5 ±0.2
79.9 ±0.7
20.1 ±0.7
84.0 ±0.5
16.0 ±0.5
GPT-4.1-mini
47.7 ±0.2
41.6 ±0.2
58.4 ±0.2
41.6 ±0.4
58.4 ±0.4
GPT-4.1-nano
28.0 ±0.2
36.4 ±0.6
63.6 ±0.6
38.2 ±0.9
61.8 ±0.9
Gemini-3.0-Flash
93.9 ±0.3
85.5 ±0.3
14.5 ±0.3
87.9 ±0.4
12.1 ±0.4
Gemini-3.1-Flash-Lite
55.4 ±1.1
18.4 ±0.9
81.6 ±0.9
30.5 ±1.5
69.5 ±1.5
Appendix
Table 7: Direct prompting results on the filtered MMLU-Pro Mathematics subset ( N=1,118 ) under the no-valid-option-aware instruction and invalid forced-choice penalty −1 . Values are mean ± standard deviation over three independent runs.
Model
Penalty
AR ↑
IFR ↓
CAR ↑
CIFR ↓
GPT-5-mini
−0.5
85.5 ±0.6
14.5 ±0.6
86.9 ±0.4
13.1 ±0.4
−1
85.6 ±0.7
14.4 ±0.7
87.2 ±0.7
12.8 ±0.7
−2
85.1 ±0.8
14.9 ±0.8
86.8 ±0.8
13.2 ±0.8
GPT-5-nano
−0.5
81.4 ±0.4
18.6 ±0.4
84.8 ±0.4
15.2 ±0.4
−1
81.1 ±1.0
18.9 ±1.0
84.5 ±1.0
15.5 ±1.0
−2
81.8 ±0.7
18.2 ±0.7
85.2 ±0.6
14.8 ±0.6
Appendix
Table 8: Full scoring-rule ablation results on the filtered MMLU-Pro Mathematics subset ( N=1,118 ) under CoT prompting and the no-valid-option-aware instruction. Only the invalid forced-choice penalty is varied.
Model
Instruction
AR ↑
IFR ↓
CAR ↑
CIFR ↓
GPT-5-mini
Penalty-only
75.2 ±0.4
24.8 ±0.4
76.5 ±0.3
23.5 ±0.3
Uncertainty-aware
76.8 ±1.1
23.2 ±1.1
78.5 ±1.1
21.5 ±1.1
No-valid-option-aware
85.6 ±0.7
14.4 ±0.7
87.2 ±0.7
12.8 ±0.7
GPT-5-nano
Penalty-only
73.4 ±0.5
26.6 ±0.5
76.6 ±0.8
23.4 ±0.8
Uncertainty-aware
76.1 ±1.3
23.9 ±1.3
79.2 ±1.2
20.8 ±1.2
No-valid-option-aware
81.1 ±1.0
18.9 ±1.0
84.5 ±1.0
15.5 ±1.0
Appendix
Table 9: Full instruction ablation results on the filtered MMLU-Pro Mathematics subset ( N=1,118 ) under CoT prompting and the +4/−1/0 scoring rule. Only the abstention instruction is varied.
Model
Parser
Acc ↑
AR ↑
IFR ↓
CAR ↑
CIFR ↓
GPT-5-mini
Robust
97.5 ±0.5
85.6 ±0.7
14.4 ±0.7
87.2 ±0.7
12.8 ±0.7
Strict
97.5 ±0.5
85.6 ±0.7
14.4 ±0.7
87.2 ±0.7
12.8 ±0.7
Stored
97.5 ±0.5
85.6 ±0.7
14.4 ±0.7
87.2 ±0.7
12.8 ±0.7
GPT-5-nano
Robust
94.1 ±0.4
81.1 ±1.0
18.9 ±1.0
84.5 ±1.0
15.5 ±1.0
Strict
94.1 ±0.4
81.1 ±1.0
18.9 ±1.0
84.5 ±1.0
15.5 ±1.0
Stored
94.1 ±0.4
81.1 ±1.0
18.9 ±1.0
84.5 ±1.0
15.5 ±1.0
Appendix
Table 10: Parser-mode sensitivity on the filtered MMLU-Pro Mathematics subset ( N=1,118 ) under the main CoT setting: no-valid-option-aware instruction and invalid forced-choice penalty −1 .
Model
Condition
Incomplete rate (%)
Invalid-format rate (%)
Eval-eligible rate (%)
GPT-5-mini
Original MCQA
0.00 ±0.00
0.33 ±0.05
99.67 ±0.05
No-valid-option
0.00 ±0.00
0.12 ±0.05
99.88 ±0.05
GPT-5-nano
Original MCQA
0.00 ±0.00
1.28 ±0.26
98.72 ±0.26
No-valid-option
0.00 ±0.00
0.15 ±0.10
99.85 ±0.10
GPT-4.1-mini
Original MCQA
0.45 ±0.09
0.21 ±0.14
99.34 ±0.19
No-valid-option
1.28 ±0.44
0.15 ±0.14
98.57 ±0.32
Appendix
Table 11: Response-quality diagnostics on the filtered MMLU-Pro Mathematics subset ( N=1,118 ) under CoT prompting and robust parsing. The no-valid-option condition uses the no-valid-option-aware instruction and invalid forced-choice penalty −1 . Values are percentages reported as mean ± standard deviation over three independent runs. Incomplete, invalid-format, and evaluation-eligible rates are all computed over the full set of 1,118 instances in each run. These three mutually exclusive response states sum to 100% before rounding. Gemini 2.5-family models below the dividing line are supplementary.
Model
Condition
Incomplete rate (%)
Invalid-format rate (%)
Eval-eligible rate (%)
GPT-5-mini
Original MCQA
0.00 ±0.00
0.20 ±0.17
99.80 ±0.17
No-valid-option
0.00 ±0.00
0.00 ±0.00
100.00 ±0.00
Qwen3-8B
Original MCQA
0.00 ±0.00
13.43 ±0.61
86.57 ±0.61
No-valid-option
0.10 ±0.17
4.41 ±1.35
95.49 ±1.22
Appendix
Table 12: Response-quality diagnostics on the filtered multi-domain MMLU subset ( N=340 ) under CoT prompting and robust parsing. The no-valid-option condition uses the no-valid-option-aware instruction and invalid forced-choice penalty −1 . Values are percentages reported as mean ± standard deviation over three independent runs. Incomplete, invalid-format, and evaluation-eligible rates are all computed over the full set of 340 instances in each run. These three mutually exclusive response states sum to 100% before rounding.
A. 1.234
F. 1.618
B. 1.414
G. 1.789
C. 2.000
H. 1.000
D. 1.581
I. 3.141
E. 2.718
J. 2.345
Appendix
Figure 4: Paired responses for the infinite-series example. In the original condition, the model selects the labeled correct option, D (1.581). After that option is removed, the model still computes 1+1/(e−1)≈1.58198 but explicitly selects the closest remaining option, E (1.618), instead of ABSTAIN .
A. 6
F. 3
B. 4
G. 2
C. 5
H. 1
D. 10
I. 9
E. 7
J. 8
Appendix
Figure 5: Paired responses for the Ramsey example. In the original condition, the model selects the labeled correct option, F (3). After that option is removed, the model still identifies R(3,2)=3 but invokes the related result R(3,3)=6 and selects A (6), instead of ABSTAIN .
Work
Answer absent by construction
Rejection expression
Abstention utility separation
Instance-level correctness conditioning
Tam et al. (2025) , None of the Above, Less of the Right
Yes
In-option NOTA
No
No
Madhusudhan et al. (2025) , Do LLMs Know When to NOT Answer?
Yes
In-option IDK/NOTA
No
No
Góral et al. (2025) , Wait, that’s not an option
Yes
Free-form rejection, detected post hoc
No
No
Wang et al. (2025) , LLMs May Perform MCQA by Selecting the Least Incorrect Option
Yes
In-option NOTA and free-form no-answer
No
Correctness across re-ordered answer options
Wang et al. (2026) , Are LLM Decisions Faithful to Verbal Confidence? (RiskEval)
No
Formal abstention action
Yes
No
Ours
Yes
Formal out-of-option ABSTAIN
Yes
Stable correctness across repeated runs
Appendix
Table 13: Comparison of related evaluation designs in terms of answer absence, rejection expression, utility separation between abstention and invalid choices, and instance-level correctness conditioning.
Multiple-choice question answering (MCQA) benchmarks in NLP use number-right scoring (accuracy), but in educational testing, the scoring scheme, the combination of the response mode models follow and the rule for grading responses, is a key design choice that dictates which abilities to reward. We examine how alternatives to number right change what MCQA measures with six education-inspired schemes that assess abilities beyond accuracy: distractor elimination, abstention, confidence calibration, and self-correction. On LLM benchmarks, these schemes: 1) shift rankings of 31 LLMs beyond rephrased number right prompts; 2) better predict the LLMs users prefer in LLM Arena; and 3) reveal distinct model capabilities, like that GPT-5 rarely abstains and readily self-corrects, while weaker open-weight models often abstain and hesitate to eliminate choices. Given the benefits of alternative scoring schemes, we discuss ways to extend them to tasks beyond MCQA.
Nishant Balepur, Paiheng Xu, Wei Ai +3
University of Maryland · New York University · Nanyang Technological University
A model should refuse two different things: answers it would get wrong, and questions it should not answer at all, such as unanswerable ones or ones resting on a false premise. The usual recipe thresholds a single confidence score, which cannot tell these apart. Across five instruction-tuned models from three families (2B to 14B), we find they are separate axes. Ordinary answer-confidence tracks whether an answer is right but is nearly blind to whether the question is answerable; a linear probe on hidden states does the reverse. The blind spot does not shrink with scale. It is worst on naturally occurring false-premise questions (CREPE). There, answer-confidence, P(IK), P(True), and even asking the model outright whether a premise is false all stay near chance, while a hidden-state probe reaches 0.69 to 0.77 AUROC: the model represents a problem it will not report. This turns out to be fixable. Instructing a model to check premises backfires, because it then disputes sound and false premises alike (57% false challenges), unable to tell them apart; routing the same instruction with the probe roughly triples challenge precision. We turn the two axes into a calibrated policy that answers only when an answerability score and a correctness score each clear a separately certifies behave differently: the unanswerable-answer rate is controllable at every scale, while the wrong-answer rate is capped by model accuracy, so the guarantee tightens as threshold policy certifies both budgets at 0.75 coverage of correct answers, against 0.31 for a single threshold; at 14B it is the only policy that certifies at all.
Language models often receive a question together with a claim about what another source answered. We audit whether such claims destabilize answers in multiple-choice question answering. For each item, we hold one wrong option fixed across misleading conditions and vary the cue template attached to it. We introduce \emph{neutral-conditioned misleading cue adoption rate} (NC-MCAR), which measures switches to that option only on valid cued trials where the same model first selected the gold answer under a neutral prompt. This is a measure of answer instability, not proof that the model knew the answer or that all deference is irrational. We evaluate four instruction-following models on MMLU-Pro and IndicMMLU-Pro in English, Hindi, Bengali, Tamil, and Telugu. Across 220{,}000 outputs, the expert template yields 41.1% aggregate NC-MCAR, compared with 12.5% for the majority template. These two conditions use the same wrong option and final instruction. Filler accuracy remains well above expert-wrong accuracy, while correct-cue prompts have high valid-response accuracy. The audit documents answer instability relevant to grounding under the tested forced-choice prompts: a bare, unverified source claim can outweigh an answer that was previously consistent with the task evidence.
Manikandan Ravikiran, Siddharth Vohra
Indian Institute of Technology Mandi, India · Carnegie Mellon University Amazon Web Services AI Native Pittsburgh, PA, USA