Mathematical reasoning benchmarks are typically organized by topic, but language models may organize their internal computation by reusable reasoning approach instead. In this paper, we investigate whether open math-capable LLMs organize internally by topical sub-skill or by reasoning approach, and we present evidence that the approach is the key. We introduce a generation-replay protocol: a model first generates a solution, after which we replay the exact prompt-plus-generation trajectory and extract activation-importance signatures over the reasoning tokens. We cluster these signatures without supervision across eight models and five mathematical reasoning sources, then evaluate the recovered structure with structural, semantic, and intervention tests. Across all 40 model-source cells, the recovered clusters outperform matched-size random baselines. Two independent frontier-LLM judges find approach-level coherence in 77-82% of real clusters versus 6-11% in within-source controls, and topic-pure clusters usually receive labels finer than the topic itself. In approach-controlled prompting, changing the requested reasoning approach shifts cluster assignment in seven of eight model conditions, whereas paraphrases largely preserve it. These results indicate that math-capable LLMs organize internal mathematical computation by reasoning approach rather than benchmark topic. The implication is that topic-stratified benchmarks and topic-balanced training corpora can still miss the axis that matters: even deliberately topic-balanced corpora may remain imbalanced over reasoning approaches.
Figures & tables
Source
Characterization
Sub-skills
Prompts
deepmind_math
Procedurally generated mathematical tasks
5
1,000
gsm8k
Grade-school arithmetic word problems
1
200
math_hendrycks
Competition-style mathematical problems
6
1,200
mathqa
Operation-annotated word problems
2
400
numinamath_1_5
Competition-level aggregated problem set
5
1,000
Total
19
3,800
Table 1: Mathematical reasoning sources used for profiling. We sample up to 200 prompts from each source/sub-skill cell and reuse the same 3,800-prompt balanced sample across all eight models.
Source
Cells
Clusters
Mean k
Mean sil.
Sil. range
deepmind_math
8
54
6.75
0.283
0.241–0.323
gsm8k
8
43
5.38
0.166
0.123–0.202
math_hendrycks
8
40
5.00
0.175
0.149–0.252
mathqa
8
42
5.25
0.200
0.158–0.249
numinamath_1_5
8
46
5.75
0.195
0.160–0.228
Total
40
225
5.62
0.204
0.123–0.323
Table 2: C1 structural result by source. All 40 cells have pemp<0.005 against 200 matched-size random partitions.
Source
Clusters
Topic-pure clusters
Topic-pure rate
gsm8k
43
43
100.0%
deepmind_math
54
28
51.9%
mathqa
42
11
26.2%
numinamath_1_5
46
1
2.2%
math_hendrycks
40
0
0.0%
All sources
225
83
36.9%
Table 3: Topic purity of recovered clusters under dominant sub-skill purity >0.70 . Topic-pure clusters are concentrated in sources where topic and approach are expected to be tightly coupled. In broad multi-approach sources, recovered clusters almost never simply recover benchmark sub-skill labels.
Claude Opus-4.7
GPT-5.4
Bundle type
n
Coherent rate
n
Coherent rate
Real clusters
225
185/225 = 82.2%
225
174/225 = 77.3%
Within-source controls
80
5/80 0 = 6.2%
80
9/80 0 = 11.2%
Gap
76.0 pp
66.1 pp
Table 4: Aggregate C2 result under the strict 4-of-5 approach-count rule. Both judges independently clear all three pre-specified thresholds: real coherent rate ≥60% , control rate ≤20% , gap ≥40 pp.
Claude Opus-4.7
GPT-5.4
Source
Real
Control
Gap
Real
Control
Gap
gsm8k
39/43 = 90.7%
1/16 = 6.2%
84.5 pp
38/43 = 88.4%
2/16 = 12.5%
75.9 pp
deepmind_math
47/54 = 87.0%
1/16 = 6.2%
80.8 pp
44/54 = 81.5%
1/16 = 6.2%
75.2 pp
mathqa
33/42 = 78.6%
1/16 = 6.2%
72.3 pp
31/42 = 73.8%
2/16 = 12.5%
61.3 pp
numinamath_1_5
35/46 = 76.1%
2/16 = 12.5%
63.6 pp
33/46 = 71.7%
3/16 = 18.8%
53.0 pp
math_hendrycks
31/40 = 77.5%
0/16 = 0.0%
77.5 pp
28/40 = 70.0%
1/16 = 6.2%
63.8 pp
Table 5: C2 by source under both judges. Both judges recover the same source-level pattern: approach-templated sources (GSM8K, DeepMind Math) produce the highest coherence rates, while broad competition-style sources (NuminaMath, Hendrycks MATH) are lower but still well above controls. Every source clears the pre-specified thresholds under both judges independently.
Model condition
C3 verdict
Shift
Paraphrases Invariance
Qwen2.5-1.5B-Instruct
confirmed
0.64
0.72
Qwen2.5-7B-Instruct
confirmed
0.64
0.84
Qwen2.5-14B-Instruct
confirmed
0.60
0.78
Ministral-8B-Instruct
confirmed
0.94
0.82
Mistral-7B-Instruct-v0.3
partial
0.96
0.48
DeepSeek-R1-Distill-Qwen-1.5B
confirmed
0.68
0.76
Table 6: C3 approach-bank results. Seven of eight model conditions are fully confirmed. Mistral-7B-Instruct passes the shift criterion but fails paraphrase invariance.
Appendix figures & tables3 assets
Supplementary material from the paper’s appendix.
Appendix
Feature space
Clusters
Purity > 0.70
Purity > 0.80
ARI
Reasoning text (TF-IDF/SVD)
225
81.8%
72.4%
0.38
Activations (generation-replay)
225
36.9%
29.8%
0.10
Appendix
Table 7: Text-based versus activation-based clustering. Both feature spaces use the same pipeline and the same generated reasoning traces across the same 40 model-source cells, all passing C1. Purity denotes dominant sub-skill purity; ARI is the adjusted Rand index against source-provided sub-skill labels. Text clusters are strongly topic-aligned; activation clusters are not. Combined with the C2 result that 77–82% of activation clusters are approach-coherent, this confirms that generation-replay signatures capture approach-level structure distinct from surface-text similarity.
Problem (abbreviated)
Sub-skill
Randy drew 5 pictures; Peter drew 3 more; Quincy drew 20 more than Peter. Total pictures? Approach: Compute each person’s count sequentially, sum.
arithmetic_word
Bill soaks clothes: 4 min per grass stain and 7 min per marinara stain. 3 grass stains, 1 marinara stain. Total time? Approach: Multiply rate × count for each category, then sum.
arithmetic_word
Jett bought a cow for 600,spent20/day on food for 40 days, and 500onvaccination.Solditfor2500. Profit? Approach: Compute each expense, sum the expenses, then subtract from revenue.
arithmetic_word
DeShawn made 12 free throws; Kayla made 50% more; Annieka made 4 fewer than Kayla. Annieka’s count? Approach: Compute each person’s count sequentially from the previous.
arithmetic_word
Kim spends 5 min on coffee, 2 min/employee on updates, and 3 min/employee on payroll. 9 employees. Total time? Approach: Compute each task’s time, then sum.
arithmetic_word
Appendix
Table 8: Five nearest-centroid exemplars from a single GSM8K cluster. All share the sub-skill label arithmetic_word , so topic labels do not differentiate this cluster from others. The activation-based clustering groups them by their shared reasoning approach: computing named intermediate quantities sequentially and then aggregating them.
Problem (abbreviated)
Sub-skill
Given x+siny=2008 and x+2008cosy=2007 with 0≤y≤π/2 , find x+y . Approach: Subtract equations to eliminate x , substitute into Pythagorean identity, factor and solve.
precalculus
Trapezoid with bases differing by 100; midsegment divides area 2:3 ; find ⌊x2/100⌋ where x bisects the area. Approach: Express areas algebraically, solve ratio equation for base length, derive equal-area segment via quadratic relation.
geometry
9x3+5ax2+4bx+a=0 has three distinct positive roots with ∑log2ri=4 ; find a . Approach: Apply Vieta’s formulas, convert logarithm sum to product, substitute into product-of-roots relation, solve.
algebra
Two noncongruent integer-sided isosceles triangles with the same perimeter and area; base ratio 8:7 ; find the minimum perimeter. Approach: Parameterize sides via ratio, equate perimeters and areas, substitute into Pythagorean height formula, solve quadratic.
geometry
Sn = sum of reciprocals of nonzero digits from 1 to 10n ; find smallest n with Sn∈Z . Approach: Count digit occurrences, express Sn=n⋅10n−1⋅7129/2520 , reduce to divisibility conditions on n .
number theory
Appendix
Table 9: Five nearest-centroid exemplars from a single cluster spanning four sub-skills. Despite covering precalculus, geometry, algebra, and number theory, all five solutions share the same reasoning approach: translate constraints into equations, eliminate variables, and solve the reduced system. A topic-based clustering would separate these problems; the activation-based clustering groups them by shared computational approach.
Mathematical reasoning is a hallmark of human intelligence, and whether large language models (LLMs) can meaningfully perform it remains a central question in artificial intelligence and cognitive science. As LLMs are increasingly integrated into scientific workflows, rigorous evaluation of their mathematical capabilities becomes a practical necessity. Existing benchmarks are limited by synthetic settings and data contamination. We present LiveMathematicianBench, a dynamic multiple-choice benchmark for research-level mathematical reasoning built from recent arXiv papers published after model training cutoffs. By grounding evaluation in newly published theorems, it provides a realistic testbed beyond memorized patterns. The benchmark introduces a thirteen-category logical taxonomy of theorem types (e.g., implication, equivalence, existence, uniqueness), enabling fine-grained evaluation across reasoning forms. It employs a proof-sketch-guided distractor pipeline that uses high-level proof strategies to construct plausible but invalid answer choices reflecting misleading proof directions, increasing sensitivity to genuine understanding over surface-level matching. We also introduce a substitution-resistant mechanism to distinguish answer recognition from substantive reasoning. Evaluation shows the benchmark is far from saturated: Gemini-3.1-pro-preview, the best model, achieves only 43.5%. Under substitution-resistant evaluation, accuracy drops sharply: GPT-5.4 scores highest at 30.6%, while Gemini-3.1-pro-preview falls to 17.6%, below the 20% random baseline. A dual-mode protocol reveals that proof-sketch access yields consistent accuracy gains, suggesting models can leverage high-level proof strategies for reasoning. Overall, LiveMathematicianBench offers a scalable, contamination-resistant testbed for studying research-level mathematical reasoning in LLMs.
Linyang He, Qiyao Yu, Hanze Dong +5
Columbia University · Microsoft Research · University of Amsterdam
Despite the pivotal role of numerical reasoning as the cornerstone of mathematical capabilities in large language models (LLMs) across applications, few benchmarks evaluate LLMs by integrating numerical processing and mathematical reasoning, hindering the interpretability of failures in math tasks. We introduce PyraMathBench, a comprehensive hierarchical benchmark with 32,505 questions derived from 7,404 math word problems, spanning 4 key cognitive aspects, 14 subcategories, and 2 modalities. Experiments reveal that LLMs' performance is severely compromised by inadequate numerical computation and weak handling of abstract numerical questions. To address this, we propose the Smart Optimization & Learning-based VErsatile module (SOLVE) and Interactive Relative Policy Optimization (IRPO), which enhance LLMs' numerical-mathematical synergy via efficient tool calls (fuzzy matching and low-quality call rejection). Comparative experiments show Qwen-2.5 achieves a 5.0 score improvement with SOLVE and IRPO training.
Zetian Ouyang, Linlin Wang, Gerard de Melo +1
East China Normal University · Hasso Plattner Institute, University of Potsdam
We present a new approach for benchmarking Large Language Model (LLM) capabilities on research-level mathematics. Existing benchmarks largely rely on static, hand-curated sets of contest or textbook-style problems as proxies for mathematical research. Instead, we establish an updatable benchmark evaluating models directly on the latest research results in mathematics. This consists of an automatic pipeline that extracts lemmas from arXiv and rewrites them into self-contained statements by making all assumptions and required definitions explicit. It results in a benchmark that can be updated regularly with new problems taken directly from human mathematical research, while previous instances can be used for training without compromising future evaluations. We benchmark current state-of-the-art LLMs, which obtain around 10-15% accuracy in theorem proving (pass@1) depending on the model, showing that there is currently a large margin of progression for LLMs to reach human-level proving capabilities in a research context.
Antoine Peyronnet, Fabian Gloeckle, Amaury Hayat
1Ecole Normale Superieure de Rennes · 2Ecole des Ponts Paris · 3Korean Institute for Advanced Study