What Makes Something Hard(er)? Explaining Question Difficulty in Natural Language
Organizations: Department of Computer Science ETH Zurich · Machine Learning and Optimization Lab EPFL
Abstract
Difficulty is one of the most fundamental properties of a question: it determines whether the question can meaningfully discriminate between models of differing ability. Although a variety of methods can now estimate or predict difficulty automatically, they yield only a single descriptive number, with no account of the underlying factors that make a question difficult in the first place. In this work, we propose a data-driven approach that automatically generates and validates natural-language hypotheses explaining what makes one question harder than another. We first estimate each item's difficulty from the responses of a large pool of LLMs using Item Response Theory. We then sample contrasting sets of easy and hard questions and prompt an LLM to propose candidate explanations of the difference, which are subsequently validated and selected on held-out questions. Experimental results across three datasets spanning mathematical, logical, and commonsense reasoning show that our method produces interpretable and predictive hypotheses. On their own, they predict the difficulty of unseen questions competitively with, or better than, advanced black-box difficulty regressors; used as additional features, they further improve those regressors, implying that they discover difficulty signals that existing models fail to capture. Moreover, we demonstrate that editing questions according to a hypothesis can shift their measured difficulty in the expected direction, indicating that the discovered hypotheses are causally valid difficulty factors rather than post-hoc descriptions. Our approach thus turns a purely descriptive difficulty score into actionable statements.
Figures & tables
| # | Top-5 Generated Hypotheses | Cohen’s | Cov. (%) | |
| GSM8K | ||||
| 1 | The problem requires multistep bookkeeping in which intermediate results feed subsequent calculations. | 0.899 ‡ | 51.0 | |
| 2 | The solver must derive a rate, ratio, or percentage rather than directly apply one. | 1.174 ‡ | 9.0 | |
| 3 | The solver must set up and solve an algebraic equation rather than perform only sequential arithmetic. | 0.790 † | 4.0 | |
| 4 | The reference answer contains a logical error or contradiction that conflicts with the correct solution. | 2.063 ‡ | 1.0 | |
| 5 | The solver must aggregate multiple time intervals or convert between time units to obtain a final duration. | 0.584 † | 6.5 | |
| Approaches | GSM8K | BBH-structured | Winogrande | |||
|---|---|---|---|---|---|---|
| RoBerta-base finetuned | 0.02 | 1.454 0.02 | 0.646 0.01 | 0.698 0.01 | -0.038 0.09 | 1.711 0.07 |
| Gemma-3-4B last-token | 0.191 0.00 | 1.637 0.00 | 0.554 0.00 | 0.783 0.00 | -0.019 0.00 | 1.697 0.00 |
| Qwen3-Embedding-4B | 0.200 0.00 | 1.629 0.00 | 0.533 0.00 | 0.802 0.00 | -0.000 0.00 | 1.681 0.00 |
| Gemini-3.1-pro 10-shot | 0.346 0.03 | 1.472 0.04 | 0.139 0.172 | 1.085 0.107 | -0.134 0.14 | 1.788 0.11 |
| Claude-sonnet-5 10-shot | 0.041 0.06 | 1.782 0.05 | 0.043 0.284 | 1.139 0.165 | 0.080 0.03 | 1.618 0.03 |
| Hypothesis | Requires tracking and calculating values for four or more distinct entities or categories. |
|---|---|
| Increase: original | Mishka bought 3 pairs of shorts, 3 pairs of pants, and 3 pairs of shoes. One pair of shorts costs 22.50 and one pair of shoes costs 243 |
| Increase: | |
| edited | Mishka bought 3 pairs of shorts, 3 pairs of pants, 3 pairs of shoes, and 3 shirts . One pair of shorts costs 22.50, one pair of shoes costs 15 . How many dollars did Mishka spend on all the clothing items? Ans. $288 |
| Decrease: original | Bernie has 4 dogs. They each need a certain amount of exercise per day. The first needs to walk 1 mile. The second needs to walk 4 miles. The third needs to walk 3 miles. On average, they need to walk 3 miles per day. How many miles does the last dog need? Ans. 4 miles |
| Decrease: edited | Bernie has 3 dogs . They each need a certain amount of exercise per day. The first needs to walk 1 mile. The second needs to walk 4 miles. On average, they need to walk 3 miles per day. How many miles does the last dog need? Ans. 4 miles |
| GSM8K | BBH-structured | |||||||
|---|---|---|---|---|---|---|---|---|
| Intervention | Orig. | Edited | Orig. | Edited | ||||
| Increase difficulty | 45.01 | 29.75 | 70.9% (100/141) | 60.63 | 36.75 | 77.6% (52/67) | ||
| Decrease difficulty | 25.61 | 57.79 | 92.2% (47/51) | 36.82 | 65.34 | 85.5% (47/55) | ||
Appendix figures & tables19 assets
Supplementary material from the paper’s appendix.
Appendix
| Split | IRT difficulty | ||||||
|---|---|---|---|---|---|---|---|
| Dataset | #Models | Min | Mean SD | Max | |||
| GSM8K | 5,000 | 719 | 400 | 200 | |||
| BBH-structured | 3,811 | 558 | 558 | 280 | |||
| Winogrande | 5,000 | 759 | 254 | 254 | |||
| Dataset | Seed | Question IDs |
|---|---|---|
| GSM8K | 0 | 3344, 3268, 2478, 3542, 3198, 2561, 2761, 3407, 3132, 2894 |
| 1 | 2823, 2876, 3727, 3336, 3078, 3406, 2555, 2866, 2806, 3139 | |
| 2 | 2946, 2602, 3689, 2950, 3273, 3149, 3272, 3306, 3005, 2880 | |
| BBH-structured | 0 | 36661, 34883, 36010, 36121, 34851, 36773, 35797, 36779, 35888, 35503 |
| 1 | 36235, 36068, 35911, 37867, 36206, 36670, 37800, 37794, 35495, 34780 | |
| 2 | 36664, 37919, 36648, 35527, 35981, 35567, 34958, 36780, 35475, 34813 |
| Increase-difficulty edits | Decrease-difficulty edits | ||||||
| # | Model | Orig. | Edited | Orig. | Edited | ||
| GSM8K | |||||||
| 1 | TinyLlama-1.1B ( Zhang et al., 2024 ) | 1.4 | 1.4 | 0.0 | 0.0 | 0.0 | 0.0 |
| 2 | Gemma-2B-IT ( Mesnard et al., 2024 ) | 4.3 | 6.4 | 2.1 | 2.0 | 25.5 | 23.5 |
| 3 | Qwen1.5-0.5B-Chat ( Bai et al., 2023 ) | 10.6 | 2.8 | 7.8 | 0.0 | 13.7 | 13.7 |
| 4 | Phi-1.5 ( Li et al., 2023 ) | 13.5 | 19.9 | 6.4 | 3.9 | 43.1 | 39.2 |
| Increase | Decrease | |||||||||
| # | Q | ✓ ✗ | ✗ ✓ | Q | ✓ ✗ | ✗ ✓ | ||||
| GSM8K | ||||||||||
| 1 | 11 | 46 | 9 | 19.7 | 75.9 | 0 | 0 | 0 | — | — |
| 2 | 11 | 48 | 5 | 22.99 | 83.4 | 0 | 0 | 0 | — | — |
| 3 | 9 | 37 | 11 | 16.99 | 68.0 | 2 | 0 | 10 | 29.41 | 29.4 |
| 4 | 28 | 97 | 40 | 11.97 | 68.3 | 5 | 2 | 25 | 27.06 | 36.5 |
| # | Hypotheses | Cohen’s | Cov. (%) | |
|---|---|---|---|---|
| GSM8K | ||||
| 1 | Questions involving counter-intuitive operations or implicit variable relationships, such as returned items or equating frequency of work to earnings, are more likely to be rated as difficulty level 3. | 1.154 | 6.0 | |
| 2 | Questions that require tracking and aggregating multiple entities or sub-totals over several steps tend to be rated as moderate to high difficulty. | 0.855 | 40.5 | |
| 3 | The inclusion of negative numbers, unit conversions, or complex ratios significantly increases the difficulty level of a question. | 0.584 | 10.5 | |
| 4 | Questions that require tracking and combining multiple sub-totals before reaching the final answer are more difficult than those with a single linear calculation path. | 0.753 | 41.0 | |
| 5 | Problems that involve implicit conversions or understanding of units, such as ’dozen’ or ’a week’, often have a higher difficulty level. | 0.396 | 11.5 | |
| # | Hypotheses | Cohen’s | Cov. (%) | |
|---|---|---|---|---|
| BBH-structured | ||||
| 1 | Questions where the final answer is explicitly stated verbatim in the prompt text are easier than those requiring implicit logical deduction. | -0.865 | 30.2 | |
| 2 | Questions that require tracking and ordering five or more entities are generally harder than those involving fewer entities, unless the answer is explicitly stated. | 0.743 | 44.0 | |
| 3 | Logical deduction tasks that require combining multiple relative constraints to deduce a complete sequence are typically rated at difficulty level 3. | 0.859 | 28.4 | |
| 4 | Questions that require sorting numerical values or combining multiple relational constraints to find an intermediate position tend to be rated at the highest difficulty level. | 1.261 | 21.5 | |
| 5 | Questions that require deducing the relative order of objects from multiple constraints are generally more difficult than questions that explicitly state the answer or involve simple counting. | 0.144 | 48.4 | |
| # | Hypotheses | Cohen’s | Cov. (%) | |
|---|---|---|---|---|
| Winogrande | ||||
| 1 | Questions requiring the model to infer physical constraints or spatial relationships, such as a bulky object breaking a band or a minuscule ring not fitting a gem, are typically harder. | 0.153 | 10.6 | |
| 2 | Increased syntactic distance or the use of complex clause structures between the target blank and the contextual clues increases the difficulty level. | — | 0.4 | |
| 3 | Questions containing explicit causal conjunctions like ’because’ or ’so’ that directly link an action to its immediate outcome are more likely to be rated as difficulty level 1. | 0.034 | 65.4 | |
| 4 | The presence of negation, sarcasm, or counter-intuitive scenarios increases the difficulty level of the question. | 0.041 | 4.7 | |
| 5 | Questions that rely on direct, common-sense word associations, such as ’chainsaw’ and ’tree cutter’, are easier (level 1) than those requiring multi-step logical deductions. | -0.550 | 7.1 | |
| # | Hypotheses | Cohen’s | Cov. (%) | |
|---|---|---|---|---|
| GSM8K | ||||
| 1 | no coherent item-level property Trigger: the final digit of the last number in an arithmetic expression, immediately preceding the equals sign. | — | — | 100.0 |
| 2 | no coherent item-level property Trigger: The space token immediately following the ’####’ marker before the final answer in GSM8K solutions. | -0.318 | 94.0 | |
| 3 | The question involves specific times of day formatted with a colon. Trigger: The colon ’:’ used in time formats (e.g., 5:00, 10:00). | 0.855 | 8.0 | |
| 4 | no coherent item-level property Trigger: A space character immediately preceding a number. | 0.612 | 39.0 | |
| 5 | no coherent item-level property Trigger: The equals sign (’ =’) immediately preceding a calculator annotation (’ ¡¡’) in the reference solution. | 0.074 | 53.0 | |
| # | Hypotheses | Cohen’s | Cov. (%) | |
|---|---|---|---|---|
| BBH-structured | ||||
| 1 | The reference solution involves calculating a sum, average, or identifying a specific numerical quantity. Trigger: The space or first digit of a number that represents the result of an arithmetic calculation or a specific numerical value in the reasoning process. | 0.453 | 31.1 | |
| 2 | The reference solution contains a list of relative ordering constraints or properties separated by commas, culminating in an ’and’. Trigger: The comma separating the penultimate and final items in a list of sequential logical deductions or conditions, typically right before ’and’. | — | — | 100.0 |
| 3 | The question asks to sort items by alphabetic order or find the second heaviest item. Trigger: The token ’abetic’ in the phrase ’alphabetic order’ or the token ’ second’ in the phrase ’The second heaviest’. | 0.745 | 3.6 | |
| 4 | The question has at least six multiple-choice options (reaching option F or beyond). Trigger: The open parenthesis ’(’ for option F, or the name ’Gwen’ in option D. | 0.450 | 27.9 | |
| 5 | The question involves logical deduction over ordered objects, specifically identifying the position of an object in a sequence. Trigger: the word ’the’ in phrases like ’from the left’ or ’from the right’ within multiple-choice options | 0.278 | 93.2 | |
| # | Hypotheses | Cohen’s | Cov. (%) | |
|---|---|---|---|---|
| Winogrande | ||||
| 1 | The reference solution explains the reasoning by attributing a state of being or property to one of the characters (e.g., being blind, Jewish, hungry, skilled, poor). Trigger: The copula verbs ’is’, ’are’, ’was’, ’be’ or the noun ’kids’ when describing a property or state of a person in the explanation of a Winograd schema. | 0.176 | 93.3 | |
| 2 | The question involves a scenario set in a workplace, business, or commercial context. Trigger: Tokens related to businesses, workplaces, or commercial entities (like ’factory’, ’plant’, ’supermarket’, ’company’, ’companies’, ’merchant’, ’office’) and their preceding determiners/modifiers. | -0.123 | 2.4 | |
| 3 | no coherent item-level property Trigger: The feature fires on nouns or names that are the object of a preposition or part of a noun phrase in the explanation section of a Winograd schema problem. | -0.227 | 2.4 | |
| 4 | The reference solution uses a complex sentence structure with relative clauses, modal verbs, or infinitive phrases to explain the reasoning. Trigger: Tokens like ’would’, ’which’, ’to’, ’that’, ’can’, or commas that introduce a relative clause, modal verb, or infinitive phrase explaining a causal or conditional relationship in the reasoning. | 0.156 | 4.3 | |
| 5 | no coherent item-level property Trigger: Tokens related to quantity, degree, or specific actions (e.g., ’any’, ’his’, ’a’, ’many’, ’less’, ’cure’, ’pose’, ’of’) in the context of reasoning explanations. | -0.012 | 2.0 | |
| # | Hypotheses | Cohen’s | Cov. (%) | |
|---|---|---|---|---|
| GSM8K | ||||
| 1 | requires calculating an average; specifically, the question asks for the mean or average of a set of values | 0.835 | 2.5 | |
| 2 | involves unit conversions; specifically, the problem requires converting between different units of measurement (e.g., minutes to hours, kb to Mb, yards to miles) | 0.430 | 8.0 | |
| 3 | contains a narrative with physical actions causing subsequent events; specifically, the text describes an action (like throwing an object) that directly causes another event (like more objects falling) | — | 0.5 | |
| 4 | involves multi-step arithmetic with more than three operations; specifically, the solution requires performing four or more distinct mathematical operations to reach the final answer | 0.954 | 46.0 | |
| 5 | uses decimal numbers; specifically, the text contains non-integer numerical values such as 1.5, 1.25, or 2.50 in the problem description | 0.025 | 9.0 | |
| # | Hypotheses | Cohen’s | Cov. (%) | |
|---|---|---|---|---|
| BBH-structured | ||||
| 1 | requires determining the position of an object relative to others; specifically, the question asks for the position of an object (e.g., ’third from the left’, ’rightmost’) based on a set of relative ordering constraints | 0.264 | 27.5 | |
| 2 | involves logical deduction of a sequence; specifically, the text requires ordering a set of objects based on a series of relative constraints | -0.041 | 54.3 | |
| 3 | requires multi-step logical deduction to determine a sequence; specifically, the text provides a set of constraints about the relative order or age of objects and asks to identify the position of a specific object | 0.114 | 49.3 | |
| 4 | includes an example sentence explaining how to read the data; specifically, the text contains a phrase like ’ | -0.018 | 6.4 | |
| 5 | has a multiple-choice format with more than five options; specifically, the text presents a question followed by six or more lettered options | 0.393 | 37.5 | |
| # | Hypotheses | Cohen’s | Cov. (%) | |
|---|---|---|---|---|
| Winogrande | ||||
| 1 | requires reasoning about inverse relationships; specifically, the text requires understanding that an increase in one quantity (e.g., eating more) implies a decrease or lack in another (e.g., having a smaller lunch or being hungry) | 0.057 | 3.9 | |
| 2 | involves negation in the premise; specifically, the text contains words like ’not’, ’didn’t’, or ’unlike’ to establish a negative condition for one of the subjects | -0.134 | 18.5 | |
| 3 | uses dimensional or spatial adjectives; specifically, the sentence containing the blank uses words like ’small’, ’big’, ’broad’, or ’confined’ to describe the missing entity | 0.178 | 10.2 | |
| 4 | requires inferring the cause of a state or action from a contrasting pair of individuals; specifically, the text contrasts two people and requires deducing which one is responsible for a specific outcome or state based on their described characteristics | -0.136 | 28.7 | |
| 5 | involves contrasting actions or states between two individuals; specifically, the text describes two people doing different things or having different characteristics, and the blank must be resolved by matching the correct person to the concluding action or state | -0.068 | 36.2 | |
| # | Hypotheses | Cohen’s | Cov. (%) | |
|---|---|---|---|---|
| GSM8K | ||||
| 1 | The question requires calculating a total by summing multiple intermediate values derived from rates, prices, or percentages | 0.108 | 26.0 | |
| 2 | The question is a math word problem that can be solved using exactly two basic arithmetic operations | -0.927 | 22.0 | |
| 3 | The question involves calculating the age of one or more individuals based on relative age differences and time shifts | 0.439 | 1.0 | |
| 4 | The question requires calculating a final value by determining and combining at least two intermediate quantities | 0.707 | 61.0 | |
| 5 | The question is a multi-step math word problem that requires accounting for implicit entities, rounding rules, or multi-period summations | 1.771 | 2.0 | |
| # | Hypotheses | Cohen’s | Cov. (%) | |
|---|---|---|---|---|
| BBH-structured | ||||
| 1 | The question asks for the position of an object whose position is explicitly stated in the provided text, requiring no actual logical deduction | -0.872 | 12.1 | |
| 2 | The question requires deducing the left-to-right order of a sequence of books based on relative position clues | -0.012 | 5.4 | |
| 3 | The question asks to identify the available time gap in a person’s schedule when they could have visited a specific location | 0.122 | 20.7 | |
| 4 | The question asks for the position of an object, and the correct answer is explicitly stated verbatim in the provided text, requiring no actual logical deduction | -0.872 | 12.1 | |
| 5 | asks for the color of the left-most item in a list of objects | -0.936 | 0.7 | |
| # | Hypotheses | Cohen’s | Cov. (%) | |
|---|---|---|---|---|
| Winogrande | ||||
| 1 | The question asks the model to choose between two options to fill in a blank or resolve a pronoun in a given sentence based on contextual clues | — | — | 100.0 |
| 2 | The question asks the user to choose the correct entity from two options to fill in a blank in a given sentence | — | — | 100.0 |
| 3 | The question requires resolving a missing word or pronoun in a sentence by choosing between two provided options | — | — | 100.0 |
| 4 | The question requires resolving an ambiguous pronoun or blank in a sentence to one of two given noun options | — | — | 100.0 |
| 5 | The question requires resolving a pronoun or blank in a sentence by choosing between two provided options, similar to a Winograd Schema Challenge | — | — | 100.0 |
| # | Hypotheses | Cohen’s | Cov. (%) | |
|---|---|---|---|---|
| GSM8K | ||||
| 1 | The question contains implicit information or requires the solver to infer a missing value from a combination of given constraints (e.g., working backwards from a total or a difference). | 0.623 | 15.0 | |
| 2 | The question involves multi-step bookkeeping where intermediate quantities must be calculated and then used as inputs for subsequent calculations, rather than a simple linear sequence of operations. | 0.899 | 51.0 | |
| 3 | The reference answer contains a logical error or contradiction, such as ignoring a stated value and summing the wrong numbers, making it difficult for a correct model to match the expected output. | 2.063 | 1.0 | |
| 4 | The question requires more than three distinct arithmetic operations to reach the final answer. | 0.408 | 26.0 | |
| 5 | Requires at least five distinct arithmetic operations to reach the final answer. | 0.648 | 20.0 | |
| # | Hypotheses | Cohen’s | Cov. (%) | |
|---|---|---|---|---|
| BBH-structured | ||||
| 1 | The question asks for a fact, property, or position that is explicitly stated in the text (e.g., ’The plums are the cheapest’, ’What color is the textbook?’), requiring only direct lookup rather than logical deduction. | -1.206 | 20.0 | |
| 2 | The question asks to identify an entity at a specific non-extreme ordinal position (e.g., ’second-to-last’, ’third from the right’) rather than at an extreme end (e.g., ’leftmost’, ’cheapest’). | 0.492 | 28.2 | |
| 3 | The question requires sorting a list of strings (e.g., names) alphabetically to determine the answer, rather than performing a simple numerical minimum/maximum operation or direct lookup. | 1.905 | 0.7 | |
| 4 | The question requires identifying an object based on its relative position to another specific object (e.g., ’directly to the right of’, ’furthest from’) rather than its absolute position in the sequence (e.g., ’left-most’, ’right-most’). | 0.432 | 6.1 | |
| 5 | The question requires reasoning about relative spatial or temporal positions (e.g., ’to the right of’, ’between what times’) rather than absolute attributes (e.g., ’is the jug mauve?’). | 0.658 | 62.1 | |
| # | Hypotheses | Cohen’s | Cov. (%) | |
|---|---|---|---|---|
| Winogrande | ||||
| 1 | The sentence structure involves a complex causal chain where the blank refers to an entity that is the cause of a state described earlier in the sentence, requiring the model to infer the cause from the effect. | 0.402 | 4.3 | |
| 2 | The easier question can be solved using strong lexical associations, common collocations, or direct causal links (e.g., ’tough stain’, ’IOS’ to ’iPhone’, ’lazy’ to ’failed’). | -0.532 | 7.5 | |
| 3 | The question requires multi-step causal reasoning to connect the premise to the resolution, where the solver must infer an unstated intermediate state or action (e.g., hating a food means eating less of it, which leaves more to harvest; or accepting a challenge implies believing one can win), rather than relying on direct semantic associations or attribute matching (e.g., ice is cold, or an animal rights activist dislikes leather). | 0.559 | 26.4 | |
| 4 | The question can be resolved using basic selectional restrictions or inherent physical properties of the entities (e.g., a cloth is soft, a video is boring, a color is bright) rather than requiring causal reasoning about the events described. | -0.524 | 11.8 | |
| 5 | The harder question involves resolving a pronoun based on an action that caused regret or a negative emotional reaction (e.g., feeling bad and vowing not to do it again). | — | — | 0.0 |