An Investigation of Robustness of LLMs in Mathematical Reasoning: Benchmarking with Mathematically-Equivalent Transformation of Advanced Mathematical Problems
Organizations: Department of Computer Science University of Illinois Urbana–Champaign Urbana, IL, USA · Department of Computer Science Stanford University Stanford, CA, USA
Abstract
Frontier large language models (LLMs) now reach near-ceiling accuracy on standard mathematical-reasoning benchmarks and gold-medal-level performance at the International Mathematical Olympiad. As these benchmarks saturate and their items leak into training data, a high score no longer shows whether a model reasons robustly or which component of its reasoning fails. To evaluate reasoning while keeping results informative and failures diagnosable, we propose GAP (Generalisation-and-Perturbation), a methodology that automatically generates mathematically equivalent variants of existing mathematics problems at scale using two disjoint, interpretable transformations: (1) surface renames, probing the binding between identifiers and latent variable roles, and (2) kernel rewrites, probing whether a high-level proof plan survives a change of mathematical setting. Compared with existing benchmarks, GAP has two key benefits: (1) novel, likely unseen variants mitigate data leakage, and (2) performance across transformation families enables failure diagnosis, each transformation testing a hypothesis about the cause of failure. We instantiate GAP on all 1,051 William Lowell Putnam Competition problems from 1938 to 2024, adding 5,255 unseen variants to form PutnamGAP, a 6,306-item competition-level mathematics corpus and the first public machine-readable dataset from the full Putnam archive. Using PutnamGAP, we evaluated 18 commercial and open-source models spanning sizes and providers. Accuracy drops across all models and variant families, most severely under kernel rewrites. This gap does not close with model strength, suggesting that even the strongest models' dominant weakness is transferring a proof plan to a changed mathematical setting, rather than handling surface changes. Further analysis provides finer failure diagnoses and potentially useful insights for improving LLM reasoning.
Figures & tables
Appendix figures & tables23 assets
Supplementary material from the paper’s appendix.
Appendix
| Subset | Pairs | Kernel decrease (pp) | Models with a decrease | Largest surface decrease (pp) |
|---|---|---|---|---|
| Judged unchanged in difficulty | 537 | 3.43 | 17/18 | 2.03 |
| Length-controlled | 333 | 5.59 | 17/18 | 3.08 |
| Variant | Valid solutions | Paired | Original (%) | Variant (%) | Decrease (pp) | Exact McNemar |
|---|---|---|---|---|---|---|
| DL | 90/100 | 86 | 59.30 | 55.81 | 3.49 | 0.607 |
| DLC | 93/100 | 86 | 58.14 | 53.49 | 4.65 | 0.481 |
| DLM | 94/100 | 87 | 58.62 | 50.57 | 8.05 | 0.189 |
| GS | 90/100 | 85 | 58.82 | 47.06 | 11.76 | 0.021 |
| KV | 79/100 | 76 | 56.58 | 32.89 | 23.68 | 0.0014 |
| Run | Original (%) | Kernel (%) | Decrease (pp) | Exact McNemar |
|---|---|---|---|---|
| 1 | 43.68 | 17.24 | 26.44 | |
| 2 | 43.68 | 21.84 | 21.84 | |
| 3 | 40.23 | 20.69 | 19.54 | |
| Mean (SD) | 22.61 (3.51) |
| Condition | Accuracy (%) | Decrease (pp) | Exact McNemar |
|---|---|---|---|
| Original | 27.5 | — | — |
| DL | 17.5 | 10.0 | 0.125 |
| DLC | 15.0 | 12.5 | 0.0625 |
| DLM | 22.5 | 5.0 | 0.625 |
| GS | 15.0 | 12.5 | 0.0625 |
| KV | 20.0 | 7.5 | 0.453 |
| Model | Original | DL ( ) | DLC ( ) | DLM ( ) | GS ( ) | KV ( ) |
| claude-opus-4 | 26.5 | 23.0 ∗ ∗∗ (–3.5) | 22.2 ∗∗ ∗ (–4.3) | 21.7 ∗∗ ∗ (–4.8) | 21.4 ∗∗ ∗ (–5.1) | 13.8 ∗∗∗ (–12.7) |
| claude-sonnet-4 | 23.0 | 20.6 ∗∗∗ (–2.5) | 19.8 ∗ ∗∗ (–3.2) | 18.6 ∗∗ ∗ (–4.4) | 18.1 ∗∗∗ (–4.9) | 11.1 ∗∗∗ (–11.9) |
| deepseek-prover | 15.5 | 15.2 ∗∗∗ (–0.3) | 14.0 ∗∗∗ (–1.5) | 12.8 ∗ ∗∗ (–2.7) | 13.7 ∗∗∗ (–1.8) | 9.2 ∗∗∗ (–6.3) |
| gemini-2.5-flash-lite | 19.8 | 18.8 ∗∗∗ (–0.9) | 16.1 ∗∗ ∗ (–3.7) | 15.8 ∗∗ ∗ (–4.0) | 15.1 ∗∗∗ (–4.7) | 6.6 ∗∗∗ (–13.2) |
| gemini-2.5-pro | 78.4 | 75.2 ∗ ∗∗ (–3.1) | 74.3 ∗∗ ∗ (–4.1) | 72.8 ∗∗∗ (–5.6) | 72.9 ∗∗∗ (–5.4) | 63.5 ∗∗∗ (–14.9) |
| gemini-2.5-flash | 42.8 | 42.6 ∗∗∗ (–0.2) | 39.0 ∗ ∗∗ (–3.8) | 40.9 ∗∗∗ (–1.9) | 37.6 ∗∗ ∗ (–5.2) | 27.6 ∗∗∗ (–15.2) |
| Stratum | DL | DLC | DLM | GS | |||
|---|---|---|---|---|---|---|---|
| Topic | |||||||
| Algebra | 32.0 | 36.5 | 36.6 | 37.1 | 35.6 | 57.3 | +21.7 |
| Analysis | 29.1 | 33.6 | 37.1 | 35.2 | 33.7 | 59.2 | +25.5 |
| Number Theory | 31.2 | 35.3 | 33.6 | 37.6 | 34.4 | 57.2 | +22.8 |
| Combinatorics | 29.1 | 32.0 | 33.7 | 34.3 | 32.3 | 55.3 | +23.1 |
| Geometry | 30.8 | 30.9 | 35.3 | 35.3 | 33.1 | 53.3 | +20.3 |
| Reference frame | Surface | KV |
| Model’s own original (renamed) | — | |
| Dataset canonical variant solution | ||
| Random-pairing noise floor | ||
| Cells direction-positive (self-anchor) | — | |
| Cells (Mann–Whitney, self-anchor) | — | |
| Cells (self-anchor) | — |
| Model | Variant | null | canonical_T2 | own_T2 |
|---|---|---|---|---|
| claude-sonnet-4 | DL | 10/30 (33%) | 16/28 (57%) | 9/30 (30%) |
| DLC | 7/30 (23%) | 13/28 (46%) | 11/30 (37%) | |
| DLM | 8/30 (27%) | 12/24 (50%) | 7/30 (23%) | |
| GS | 8/30 (27%) | 8/28 (29%) | 5/30 (17%) | |
| KV | 3/30 (10%) | 4/29 (14%) | — | |
| gemini-2.5-flash | DL | 8/30 (27%) | 14/30 (47%) | 9/30 (30%) |
| Variant | null | canonical_T2 | own_T2 |
|---|---|---|---|
| DL | 26.7 [19.1, 35.8] | 44.1 [34.9, 53.8] | 31.4 [23.3, 40.8] |
| DLC | 22.8 [16.1, 31.3] | 41.3 [32.4, 51.0] | 30.7 [23.0, 39.7] |
| DLM | 24.8 [17.7, 33.5] | 37.6 [28.8, 47.4] | 30.1 [22.4, 39.1] |
| GS | 23.6 [16.7, 32.4] | 35.3 [26.7, 44.9] | 22.7 [15.9, 31.4] |
| KV | 15.8 [10.4, 23.4] | 23.3 [16.5, 31.7] | — |
| Contrast (A B) | OR | McNemar | |||
|---|---|---|---|---|---|
| canonical_T2 null | 90 | 28 | 118 | ||
| own_T2 null | 63 | 44 | 1.43 | 107 | |
| canonical_T2 own_T2 | 75 | 36 | 2.08 | 111 | |
| KV : canonical_T2 null | 15 | 6 | 2.50 | 21 |
| Model | DL | DLC | DLM | GS |
|---|---|---|---|---|
| claude-opus-4 | 2 | 1 | 63 | 26 |
| claude-sonnet-4 | 1 | 1 | 4 | 3 |
| deepseek-prover | 0 | 2 | 0 | 2 |
| gemini-2.5-flash-lite | 5 | 10 | 13 | 7 |
| gemini-2.5-pro | 5 | 8 | 10 | 6 |
| gemini-2.5-flash | 4 | 6 | 4 | 6 |
| Provider pair | Cohen’s |
|---|---|
| GPT-4o vs Claude-Sonnet-4 | 1.00 |
| GPT-4o vs Gemini-2.5-flash | 0.96 |
| Claude-Sonnet-4 vs Gemini-2.5-flash | 0.96 |
| Outcome | Candidates | Fraction |
|---|---|---|
| Accepted without repair | 987 | 93.9% |
| Entered repair at least once | 64 | 6.1% |
| Discarded after the full loop | 0 | 0.0% |
| Volume (years) | Reference |
|---|---|
| I (1938–1964) | Gleason et al. (1980) |
| II (1965–1984) | Alexanderson et al. (1985) |
| III (1985–2000) | Kedlaya et al. (2002) |
| IV (2001–2016) | Kedlaya et al. (2020) |
| PutnamGAP source subset | Sources | Surface decrease (pp) | Kernel decrease (pp) |
|---|---|---|---|
| Covered by Putnam-AXIOM | 518 | 2.71 | 12.08 |
| Newly covered | 533 | 1.75 | 7.67 |
| Prompt variant | Accuracy (%) | 95% CI | vs. base |
|---|---|---|---|
| Base solving prompt | 48 | [0.385, 0.577] | — |
| + short canonicalization hint | 58 | [0.482, 0.672] | 0.077 |
| + long canonicalization with “Rename summary” | 53 | [0.433, 0.625] | 0.441 |
| Yang et al. (2025) | GAP | |
|---|---|---|
| Controlled change | Numerical values in a fixed GSM8K template | Identifier realisation, or mathematical setting |
| Diagnostic split | Calculation vs. reasoning error in the output | Role-binding failure vs. proof-plan-transfer failure |
| When the split is defined | After observing an incorrect output | Before evaluation, by the transformation family |
| Evaluation scope | 250 GSM8K sources, 6 open models | 1,051 Putnam sources, 18 commercial and open models |