An Investigation of Robustness of LLMs in Mathematical Reasoning: Benchmarking with Mathematically-Equivalent Transformation of Advanced Mathematical Problems
Authors: Yuren Hao, Xiang Wan, ChengXiang Zhai
Organizations: Department of Computer Science University of Illinois Urbana–Champaign Urbana, IL, USA · Department of Computer Science Stanford University Stanford, CA, USA
Frontier large language models (LLMs) now reach near-ceiling accuracy on standard mathematical-reasoning benchmarks and gold-medal-level performance at the International Mathematical Olympiad. As these benchmarks saturate and their items leak into training data, a high score no longer shows whether a model reasons robustly or which component of its reasoning fails. To evaluate reasoning while keeping results informative and failures diagnosable, we propose GAP (Generalisation-and-Perturbation), a methodology that automatically generates mathematically equivalent variants of existing mathematics problems at scale using two disjoint, interpretable transformations: (1) surface renames, probing the binding between identifiers and latent variable roles, and (2) kernel rewrites, probing whether a high-level proof plan survives a change of mathematical setting. Compared with existing benchmarks, GAP has two key benefits: (1) novel, likely unseen variants mitigate data leakage, and (2) performance across transformation families enables failure diagnosis, each transformation testing a hypothesis about the cause of failure. We instantiate GAP on all 1,051 William Lowell Putnam Competition problems from 1938 to 2024, adding 5,255 unseen variants to form PutnamGAP, a 6,306-item competition-level mathematics corpus and the first public machine-readable dataset from the full Putnam archive. Using PutnamGAP, we evaluated 18 commercial and open-source models spanning sizes and providers. Accuracy drops across all models and variant families, most severely under kernel rewrites. This gap does not close with model strength, suggesting that even the strongest models' dominant weakness is transferring a proof plan to a changed mathematical setting, rather than handling surface changes. Further analysis provides finer failure diagnoses and potentially useful insights for improving LLM reasoning.
Figures & tables
Figure 1: One Putnam item ( 1938-A-2 , the buoy problem) under each of GAP’s five variants. The four surface families share the problem’s algebra and reference proof; only the names of the free variables and fixed parameters change. The kernel variant changes the setting (number of cones, painted-area convention) and therefore the answer, but follows the same five-step proof plan as the original. Surface variants stress representation ; the kernel variant stresses plan persistence .
Figure 2: Kernel-variant pipeline, illustrated on arithmetic–geometric mean inequality. The proof DAG built from the canonical solution (stage 1) is abstracted into method labels (stage 2); a replacement is generated at the leaf (stage 3) and propagated through the DAG (stage 4); the modified DAG is rendered back into a question statement (stage 5). The plan is preserved by construction, so kernel variants stress plan persistence under setting change rather than local algebra (Appendix E.3 ).
Figure 3: Accuracy by variant for the 18 evaluated models, sorted by performance on the original split. Each model’s six variant scores are connected by a thin segment so the per-model trajectory is visible at a glance. Two patterns are nearly universal: (i) the surface variants form a tight cluster, with DLC , DLM , and GS all sitting marginally below the original and DL essentially co-located with it; (ii) the kernel variant lies clearly to the left of the surface cluster for every model with non-trivial baseline accuracy, and the gap between the surface cluster and the kernel point remains large even for the strongest models. This pattern is the empirical phenomenon that Section 5 dissects.
Figure 4: Structural-overlap Cohen’s d (stable vs brittle) under two reference frames. Each column is one model; each row is one surface variant. The self-anchor effect (top) collapses under canonical anchor (bottom).
Appendix figures & tables23 assets
Supplementary material from the paper’s appendix.
Appendix
Subset
Pairs
Kernel decrease (pp)
Models with a decrease
Largest surface decrease (pp)
Judged unchanged in difficulty
537
3.43 [2.12,4.79]
17/18
2.03
Length-controlled
333
5.59
17/18
3.08
Appendix
Table 1: Kernel accuracy decrease on two difficulty-controlled subsets, macro-averaged over the 18 evaluated models. The last column gives the largest mean decrease among the four surface families on the same source items.
Variant
Valid solutions
Paired N
Original (%)
Variant (%)
Decrease (pp)
Exact McNemar p
DL
90/100
86
59.30
55.81
3.49
0.607
DLC
93/100
86
58.14
53.49
4.65
0.481
DLM
94/100
87
58.62
50.57
8.05
0.189
GS
90/100
85
58.82
47.06
11.76
0.021
KV
79/100
76
56.58
32.89
23.68
0.0014
Appendix
Table 2: DeepSeek-V4-Pro on the frozen 100-source sample. Original accuracy is computed on the paired subset for each variant, so it differs slightly across rows.
Run
Original (%)
Kernel (%)
Decrease (pp)
Exact McNemar p
1
43.68
17.24
26.44
1.55×10−6
2
43.68
21.84
21.84
1.57×10−4
3
40.23
20.69
19.54
4.88×10−4
Mean (SD)
22.61 (3.51)
Appendix
Table 3: Three independent DeepSeek-V4-Flash runs on the same 87 sources.
Condition
Accuracy (%)
Decrease (pp)
Exact McNemar p
Original
27.5
—
—
DL
17.5
10.0
0.125
DLC
15.0
12.5
0.0625
DLM
22.5
5.0
0.625
GS
15.0
12.5
0.0625
KV
20.0
7.5
0.453
Appendix
Table 4: Intel DeepMath agent on 40 sources from the frozen sample (exploratory).
Model
Original
DL ( Δ )
DLC ( Δ )
DLM ( Δ )
GS ( Δ )
KV ( Δ )
claude-opus-4
26.5
23.0 ∗ ∗∗ (–3.5)
22.2 ∗∗ ∗ (–4.3)
21.7 ∗∗ ∗ (–4.8)
21.4 ∗∗ ∗ (–5.1)
13.8 ∗∗∗ (–12.7)
claude-sonnet-4
23.0
20.6 ∗∗∗ (–2.5)
19.8 ∗ ∗∗ (–3.2)
18.6 ∗∗ ∗ (–4.4)
18.1 ∗∗∗ (–4.9)
11.1 ∗∗∗ (–11.9)
deepseek-prover
15.5
15.2 ∗∗∗ (–0.3)
14.0 ∗∗∗ (–1.5)
12.8 ∗ ∗∗ (–2.7)
13.7 ∗∗∗ (–1.8)
9.2 ∗∗∗ (–6.3)
gemini-2.5-flash-lite
19.8
18.8 ∗∗∗ (–0.9)
16.1 ∗∗ ∗ (–3.7)
15.8 ∗∗ ∗ (–4.0)
15.1 ∗∗∗ (–4.7)
6.6 ∗∗∗ (–13.2)
gemini-2.5-pro
78.4
75.2 ∗ ∗∗ (–3.1)
74.3 ∗∗ ∗ (–4.1)
72.8 ∗∗∗ (–5.6)
72.9 ∗∗∗ (–5.4)
63.5 ∗∗∗ (–14.9)
gemini-2.5-flash
42.8
42.6 ∗∗∗ (–0.2)
39.0 ∗ ∗∗ (–3.8)
40.9 ∗∗∗ (–1.9)
37.6 ∗∗ ∗ (–5.2)
27.6 ∗∗∗ (–15.2)
Appendix
Table 5: Model accuracy across families (percent scale). Significance markers from paired two-proportion z -test against the original split. McNemar contrasts on the same pairs are tabulated in aggregations/mcnemar_results.csv .
Stratum
DL
DLC
DLM
GS
Lsurf
Lkern
Δ
Topic
Algebra
32.0
36.5
36.6
37.1
35.6
57.3
+21.7
Analysis
29.1
33.6
37.1
35.2
33.7
59.2
+25.5
Number Theory
31.2
35.3
33.6
37.6
34.4
57.2
+22.8
Combinatorics
29.1
32.0
33.7
34.3
32.3
55.3
+23.1
Geometry
30.8
30.9
35.3
35.3
33.1
53.3
+20.3
Appendix
Table 6: Capability-loss rate L=b/(b+s) by topic, difficulty, and form, expressed as percentages and pooled across the 18 evaluated models. Lsurf is the simple mean of the four surface-family rates ( DL , DLC , DLM , GS ); Δ is Lkern−Lsurf . Computed from aggregations/all_results_summary.csv .
Figure 5: Positioning of GAP at the intersection of three desiderata: end-to-end automation with equivalence guarantees, competition-level difficulty, and perturbation-axis decomposition for mechanism analysis. No prior benchmark occupies all three. Only RRB ( Golikov et al., 2026 ) , which appeared after our work, occupies two of the three.
Figure 6: Accuracy on each variant family broken out by topic category.
Figure 7: Accuracy on each variant family broken out by problem form (proof vs. calculation).
Figure 8: Per-model paired accuracy delta between original and each variant family, sorted by aggregate drop. Kernel rewrites occupy the rightmost (largest-drop) end for almost every model.
Reference frame
Surface
KV
Model’s own original (renamed)
+0.75
—
Dataset canonical variant solution
+0.23
+0.35
Random-pairing noise floor
≈0
≈0
Cells direction-positive (self-anchor)
72/72
—
Cells p<0.05 (Mann–Whitney, self-anchor)
69/72
—
Cells p<10−6 (self-anchor)
40/72
—
Appendix
Table 7: Stable-vs-brittle Cohen’s d for token Jaccard under three reference frames, averaged across cells. Surface = mean over 18 models × 4 variants (72 cells); KV = mean over 18 models. The self-anchor effect ( d=+0.75 ) collapses under canonical anchor ( d=+0.23 ) on surface, demonstrating that brittle surface failures are characterised by drift from the model’s own representation, not from the canonical right answer. Kernel failures show the opposite pattern.
Model
Variant
null
canonical_T2
own_T2
claude-sonnet-4
DL
10/30 (33%)
16/28 (57%)
9/30 (30%)
DLC
7/30 (23%)
13/28 (46%)
11/30 (37%)
DLM
8/30 (27%)
12/24 (50%)
7/30 (23%)
GS
8/30 (27%)
8/28 (29%)
5/30 (17%)
KV
3/30 (10%)
4/29 (14%)
—
gemini-2.5-flash
DL
8/30 (27%)
14/30 (47%)
9/30 (30%)
Appendix
Table 8: Per-cell rebound rates ( k/n ) under three prefix conditions across 4 representative models and 5 variants. KV has only two conditions because the model’s own original trajectory does not apply (the kernel variant changes the mathematics).
Variant
null
canonical_T2
own_T2
DL
26.7 [19.1, 35.8]
44.1 [34.9, 53.8]
31.4 [23.3, 40.8]
DLC
22.8 [16.1, 31.3]
41.3 [32.4, 51.0]
30.7 [23.0, 39.7]
DLM
24.8 [17.7, 33.5]
37.6 [28.8, 47.4]
30.1 [22.4, 39.1]
GS
23.6 [16.7, 32.4]
35.3 [26.7, 44.9]
22.7 [15.9, 31.4]
KV
15.8 [10.4, 23.4]
23.3 [16.5, 31.7]
—
Appendix
Table 9: Pooled rebound rates with 95% Wilson confidence intervals across the 4 rescue models. Surface variants gain +12 to +19 pp from canonical_T2 ; KV gains only +7.4 pp.
Contrast (A > B)
b
c
OR
McNemar p
ndisc
canonical_T2 > null
90
28
3.21
8.9×10−9
118
own_T2 > null
63
44
1.43
8.1×10−2
107
canonical_T2 > own_T2
75
36
2.08
2.7×10−4
111
KV : canonical_T2 > null
15
6
2.50
7.8×10−2
21
Appendix
Table 10: Pooled paired McNemar contrasts across all 4 surface variants × 4 models. canonical_T2 significantly outperforms both null and own_T2 ; own_T2 does not significantly outperform null . The KV contrast is much weaker.
Figure 9: Rebound rate by variant family and prefix condition, pooled across the 4 rescue models ( n≈100 – 120 per cell, 95% Wilson CI). canonical_T2 dominates on every surface variant; the kernel-variant rebound is consistently smaller than the surface rebound.
Figure 10: Per-cell rescue rates: own_T2 vs canonical_T2 for each (model, surface variant). Most cells lie below the diagonal (canonical wins). The four gpt-4o-mini cells are the only points above the diagonal, indicating that for the weakest model the model’s own renamed prefix outperforms expert canonical prose.
Model
DL
DLC
DLM
GS
claude-opus-4
2
1
63
26
claude-sonnet-4
1
1
4
3
deepseek-prover
0
2
0
2
gemini-2.5-flash-lite
5
10
13
7
gemini-2.5-pro
5
8
10
6
gemini-2.5-flash
4
6
4
6
Appendix
Table 11: Collapse rate (collapse / total brittle, percentage) per (model, variant) cell. “—” indicates a cell with insufficient flip cases.
Provider pair
Cohen’s κ
GPT-4o vs Claude-Sonnet-4
1.00
GPT-4o vs Gemini-2.5-flash
0.96
Claude-Sonnet-4 vs Gemini-2.5-flash
0.96
Appendix
Table 12: Three-provider grader cross-check Cohen’s κ on 50 stratified mixed solutions. All three providers agree on 49 – 50 of 50 cases.
Outcome
Candidates
Fraction
Accepted without repair
987
93.9%
Entered repair at least once
64
6.1%
Discarded after the full loop
0
0.0%
Appendix
Table 13: Outcomes of the kernel repair-and-verify loop over the full corpus.
Volume (years)
Reference
I (1938–1964)
Gleason et al. (1980)
II (1965–1984)
Alexanderson et al. (1985)
III (1985–2000)
Kedlaya et al. (2002)
IV (2001–2016)
Kedlaya et al. (2020)
Appendix
Table 14: Primary sources for PutnamGAP. All four monographs are published by MAA Press and distributed by AMS.
PutnamGAP source subset
Sources
Surface decrease (pp)
Kernel decrease (pp)
Covered by Putnam-AXIOM
518
2.71
12.08
Newly covered
533
1.75
7.67
Appendix
Table 15: Accuracy decrease by archive coverage, averaged over the 18 evaluated models. The surface column is the mean over the four surface families.
Prompt variant
Accuracy (%)
95% CI
p vs. base
Base solving prompt
48
[0.385, 0.577]
—
+ short canonicalization hint
58
[0.482, 0.672]
0.077
+ long canonicalization with “Rename summary”
53
[0.433, 0.625]
0.441
Appendix
Table 16: Canonicalization-prompt pilot. o3 on a 100-example GS subset under three solver-prompt variants; p is the McNemar test against the base prompt.
Yang et al. (2025)
GAP
Controlled change
Numerical values in a fixed GSM8K template
Identifier realisation, or mathematical setting
Diagnostic split
Calculation vs. reasoning error in the output
Role-binding failure vs. proof-plan-transfer failure
When the split is defined
After observing an incorrect output
Before evaluation, by the transformation family
Evaluation scope
250 GSM8K sources, 6 open models
1,051 Putnam sources, 18 commercial and open models
Appendix
Table 17: Two complementary decompositions of mathematical-reasoning failures.
Mathematical reasoning is a hallmark of human intelligence, and whether large language models (LLMs) can meaningfully perform it remains a central question in artificial intelligence and cognitive science. As LLMs are increasingly integrated into scientific workflows, rigorous evaluation of their mathematical capabilities becomes a practical necessity. Existing benchmarks are limited by synthetic settings and data contamination. We present LiveMathematicianBench, a dynamic multiple-choice benchmark for research-level mathematical reasoning built from recent arXiv papers published after model training cutoffs. By grounding evaluation in newly published theorems, it provides a realistic testbed beyond memorized patterns. The benchmark introduces a thirteen-category logical taxonomy of theorem types (e.g., implication, equivalence, existence, uniqueness), enabling fine-grained evaluation across reasoning forms. It employs a proof-sketch-guided distractor pipeline that uses high-level proof strategies to construct plausible but invalid answer choices reflecting misleading proof directions, increasing sensitivity to genuine understanding over surface-level matching. We also introduce a substitution-resistant mechanism to distinguish answer recognition from substantive reasoning. Evaluation shows the benchmark is far from saturated: Gemini-3.1-pro-preview, the best model, achieves only 43.5%. Under substitution-resistant evaluation, accuracy drops sharply: GPT-5.4 scores highest at 30.6%, while Gemini-3.1-pro-preview falls to 17.6%, below the 20% random baseline. A dual-mode protocol reveals that proof-sketch access yields consistent accuracy gains, suggesting models can leverage high-level proof strategies for reasoning. Overall, LiveMathematicianBench offers a scalable, contamination-resistant testbed for studying research-level mathematical reasoning in LLMs.
Linyang He, Qiyao Yu, Hanze Dong +5
Columbia University · Microsoft Research · University of Amsterdam
Large Language Models (LLMs) achieve impressive accuracy on mathematical reasoning benchmarks, yet their performance drops when problems are modified with simple changes like different names or numbers. Code execution methods, which let models generate and run Python code instead of reasoning in natural language, have been proposed as a solution, but their effect on reasoning robustness (the ability to maintain accuracy across problem variations) has not been systematically tested. This study evaluates three approaches on 1,000 problems from the GSM-Symbolic dataset: pure reasoning using chain-of-thought (CoT) prompting, single-shot code execution using Program-Aided Language models (PAL), and iterative code execution using Step-by-Step Coding (SBSC). All three were run on paired original and modified problems using Claude Haiku 4.5. CoT was the most robust method, with an accuracy drop of 1.3 percentage points and 1.8% of problems breaking under perturbation. PAL was the least robust at 1.7 percentage points and 3.1% broke, with SBSC falling in between. Although these differences were not statistically significant (p=.096), the directional trend was consistent across all measures, suggesting that code execution, whether single-shot or iterative, does not improve reasoning robustness on grade-school-level problem variations.
Large language models (LLMs) are becoming increasingly capable mathematical collaborators, but static benchmarks are no longer sufficient for evaluating progress: they are often narrow in scope, quickly saturated, and rarely updated. This makes it hard to compare models reliably and track progress over time. Instead, we need evaluation platforms: continuously maintained systems that run, aggregate, and analyze evaluations across many benchmarks to give a comprehensive picture of model performance within a broad domain. In this work, we build on the original MathArena benchmark by substantially broadening its scope from final-answer olympiad problems to a continuously maintained evaluation platform for mathematical reasoning with LLMs. MathArena now covers a much wider range of tasks, including proof-based competitions, research-level arXiv problems, and formal proof generation in Lean. Additionally, we maintain a clear evaluation protocol for all models and regularly design new benchmarks as model capabilities improve to ensure that MathArena remains challenging. Notably, the strongest model, GPT-5.5, now reaches 98% on the 2026 USA Math Olympiad and 74% on research-level questions, showing that frontier models can now comfortably solve extremely challenging mathematical problems. This highlights the importance of continuously maintained evaluation platforms like MathArena to track the rapid progress of LLMs in mathematical reasoning.
Jasper Dekoninck, Nikola Jovanović, Tim Gehrunger +4
ETH Zurich · INSAIT · Sofia University "St. Kliment Ohridski"