Large language models (LLMs) should abstain from scientific multiple-choice questions when no option is valid, but frequent abstention alone does not demonstrate sensitivity to answer availability. We introduce ArcticQA, a dataset of 194 questions derived from primary Arctic research, with automated checks of answer support and distractor contradiction against source evidence. We further develop ArcticAbstain, a paired benchmark comparing answer-present and answer-absent conditions, with the correct answer replaced by a distractor in the latter and an explicit abstention option in both. We evaluate eight models from the Gemini, Claude, and ChatGPT families at high reasoning effort, with three trials per condition, yielding 9,312 recorded responses. Answer-present abstention rates range from 0.0% to 63.0%, whereas replacing the correct answer increases abstention by 5.05 percentage points on average. These findings highlight substantial baseline differences and the need to evaluate abstention frequency and responsiveness jointly. The dataset and benchmark are available at https://github.com/BenWilcox8/arctic-qa.
Figures & tables
Model
Present abst.
Absent abst.
Shift [95% CI] (pp)
Permutation p
Holm p
Gemini 3.8 Flash
6.4%
10.3%
+3.9[0.6,7.4]
0.037
0.187
Gemini 3.7 Flash
6.0%
8.6%
+2.6[−0.1,5.4]
0.063
0.198
Claude Fable 5.1
36.1%
46.9%
+10.8[5.7,16.1]
0.00010
0.0008
Claude Opus 5
41.3%
47.5%
+6.2[0.9,11.6]
0.031
0.187
Claude Sonnet 5
48.5%
51.0%
+2.5[−1.7,6.9]
0.331
0.331
ChatGPT Astra
63.0%
74.1%
+11.1[5.7,16.6]
0.00015
0.0011
Table 1: Condition-specific abstention rates pooled across three trials for 194 questions. Shifts are answer-absent minus answer-present rates, in percentage points (pp). The 95% confidence intervals use question-level bootstrap resampling. Permutation tests use question-level paired differences; Holm-adjusted p -values account for the eight model-specific tests.
Figure 1: Abstention rates under the answer-present and answer-absent conditions for the eight models. Each segment connects a model’s two rates, pooled across three trials. Labels report the answer-absent minus answer-present shift in percentage points.
Model
ACC
Abst. rate
F1abs
R-Acc
Gemini 3.8 Flash
0.329
0.084
0.134
0.303
Gemini 3.7 Flash
0.326
0.073
0.113
0.306
Claude Fable 5.1
0.458
0.415
0.465
0.379
Claude Opus 5
0.438
0.444
0.459
0.359
Claude Sonnet 5
0.408
0.498
0.464
0.303
ChatGPT Astra
0.541
0.686
0.618
0.540
Table 2: Metrics pooled across the three trials for eight models evaluated at high reasoning effort on 194 Arctic scientific questions. Each trial’s metrics combine the answer-present and answer-absent conditions.
Multiple-choice question answering (MCQA) is commonly used to evaluate large language models under the assumption that one of the provided options is correct, typically using answer-selection accuracy. However, in real deployments, users or retrieval systems may provide invalid option sets in which none of the listed choices is correct, and selecting one of them may incur downstream cost. We study this setting as penalty-framed no-valid-option MCQA. Using the mathematics subset of MMLU-Pro, we remove the labeled correct option, allow models to either choose a remaining option or output ABSTAIN, and penalize invalid forced-choice responses. We further introduce correct-conditioned analysis, evaluating abstention only on instances that the model originally answered correctly. Experiments show that high MCQA accuracy does not fully guarantee abstention reliability: even under explicit no-valid-option-aware instructions and penalty-based scoring, models still produce invalid forced-choice responses for a subset of originally correct instances. These results show that penalty-framed no-valid-option MCQA reveals an aspect of model reliability not captured by standard answer-selection accuracy.
Jinhyeok Kim, Hye-Young Jung
Applied Artificial Intelligence Hanyang University, Republic of Korea · Mathematical Data Science Hanyang University, Republic of Korea
Large language models can produce fluent answers when their factual support is weak. This paper introduces Chain-of-Self-Questioning (CoSQ), a prompt-only framework that makes answer commitment conditional on an explicit assessment of the information required to answer a question. We evaluate three CoSQ variants under seventeen conditions on the 817-item TruthfulQA multiple-choice validation set using eleven open-weight and hosted model families. In the final balanced-option protocol, Grounded-CoSQ at τ=0.90 reduces the mean unconditional wrong-commitment rate from 13.1% under chain-of-thought prompting to 8.9%, a 32.1% relative reduction, while increasing answered accuracy from 86.9% to 89.7% and answering 87.6% of questions. Both improvements hold for all eleven models and at every evaluated threshold. Critical-CoSQ and Adaptive-CoSQ provide neighboring operating points with 88.6% and 86.5% coverage, respectively, while remaining more reliable than the baseline. A secondary Natural Questions Short-Answer evaluation provides convergent open-form evidence. These findings show that self-assessment can support explicit, tunable answer-or-abstain decisions when an unsupported commitment is more costly than referral or review.
Ali Şenol
Department of Computer Engineering, Tarsus University, Tarsus, Türkiye
Reliable evaluation of large language models should separate supported answering from unsupported guessing without conflating either with data contamination, prompt idiosyncrasy, or generic refusal behavior. We present a contamination-aware, multi-zone benchmark for measuring the transition from answerable knowledge to abstention-expected unknowns under frozen build-time labels. The benchmark contains 1,200 items across five domains, explicit abstention expectations, contamination-risk metadata, and dual parsing with an official strict parser plus a normalized robustness parser. We evaluate FLAN-T5, Qwen2.5-Instruct, and Llama-3-Instruct models under locked answer-or-abstain prompts, answer-only controls, and prompt-template variants. The benchmark is not solved by generic non-answer behavior: FLAN baselines remain weak on productive abstention, while stronger instruction-tuned models expose a selective but incomplete transition from answering to abstaining. Qwen2.5-3B-Instruct achieves the best overall reliability, but answer-expected zones remain difficult, calibration remains poor, and benign-item refusal persists. Prompt and parser robustness analyses preserve the main ranking and qualitative conclusions. The benchmark therefore provides a reproducible protocol for auditing answerability, abstention, refusal, and contamination as distinct but interacting dimensions of LLM reliability.The dataset is publicly available at https://github.com/renweimeng/Know2Guess-A-Contamination-Aware-Multi-Zone-Benchmark.