cs.CLOct 7, 2026

Arctic Questions, Missing Answers: A Dataset and Benchmark for LLM Abstention in Arctic Science

Authors: Benjamin Wilcox, Dawei Gao, Pradeeban Kathiravelu, Douglas Causey, Kewei Sha, Yunhe Feng

Organizations: University of North Texas, Denton, TX · University of Alaska Anchorage, Anchorage, AK

Abstract

Large language models (LLMs) should abstain from scientific multiple-choice questions when no option is valid, but frequent abstention alone does not demonstrate sensitivity to answer availability. We introduce ArcticQA, a dataset of 194 questions derived from primary Arctic research, with automated checks of answer support and distractor contradiction against source evidence. We further develop ArcticAbstain, a paired benchmark comparing answer-present and answer-absent conditions, with the correct answer replaced by a distractor in the latter and an explicit abstention option in both. We evaluate eight models from the Gemini, Claude, and ChatGPT families at high reasoning effort, with three trials per condition, yielding 9,312 recorded responses. Answer-present abstention rates range from 0.0% to 63.0%, whereas replacing the correct answer increases abstention by 5.05 percentage points on average. These findings highlight substantial baseline differences and the need to evaluate abstention frequency and responsiveness jointly. The dataset and benchmark are available at https://github.com/BenWilcox8/arctic-qa.

Figures & tables

Explore similar work

CardsList
  1. Penalty-Framed No-Valid-Option MCQA: Analyzing LLM Abstention under Invalid Choices

    Oct 6, 2026Jinhyeok Kim, Hye-Young JungMultiple-Choice Questions

  2. When Should LLMs Abstain? Chain-of-Self-Questioning for Selective Risk Control

    Sep 15, 2026Ali ŞenolAbstentionCommitment