cs.CLOct 1, 2026

What Makes Something Hard(er)? Explaining Question Difficulty in Natural Language

Authors: Peng Cui, Qiaoyuan Zheng, Rudolf Debelak, Mrinmaya Sachan

Organizations: Department of Computer Science ETH Zurich · Machine Learning and Optimization Lab EPFL

Abstract

Difficulty is one of the most fundamental properties of a question: it determines whether the question can meaningfully discriminate between models of differing ability. Although a variety of methods can now estimate or predict difficulty automatically, they yield only a single descriptive number, with no account of the underlying factors that make a question difficult in the first place. In this work, we propose a data-driven approach that automatically generates and validates natural-language hypotheses explaining what makes one question harder than another. We first estimate each item's difficulty from the responses of a large pool of LLMs using Item Response Theory. We then sample contrasting sets of easy and hard questions and prompt an LLM to propose candidate explanations of the difference, which are subsequently validated and selected on held-out questions. Experimental results across three datasets spanning mathematical, logical, and commonsense reasoning show that our method produces interpretable and predictive hypotheses. On their own, they predict the difficulty of unseen questions competitively with, or better than, advanced black-box difficulty regressors; used as additional features, they further improve those regressors, implying that they discover difficulty signals that existing models fail to capture. Moreover, we demonstrate that editing questions according to a hypothesis can shift their measured difficulty in the expected direction, indicating that the discovered hypotheses are causally valid difficulty factors rather than post-hoc descriptions. Our approach thus turns a purely descriptive difficulty score into actionable statements.

Figures & tables

Appendix figures & tables19 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Question Difficulty Estimation for Large Language Models via Answer Plausibility Scoring

    May 12, 2026Jamshid Mozafari, Bhawna Piryani, Adam JatowtPhysical Plausibility

  2. Cognitive Episodes in LLM Reasoning Traces Enable Interpretable Human Item Difficulty Prediction

    Jun 26, 2026Chenguang Wang, Ming Li, Xinyue Zeng +4LLM Reasoning StrategiesLarge Reasoning Models

  3. Estimating Item Difficulty with Large Language Models as Experts

    May 18, 2026Diana Kolesnikova, Kirill Fedyanin, Abe D. Hofman +2Item Response TheoryRater