Technical-domain benchmarks constructed from arXiv papers can overlap the same public text used in LLM pretraining. Auditing this risk at training-corpus scale requires searching billions of corpus n-grams while retaining per-item evidence that users can inspect and recompute. We contribute a reverse-probe audit at a fixed 13-gram threshold: it indexes benchmark prompts, streams public pretraining corpora, and emits per-problem prompt-surface lexical-overlap metadata with memory that scales with the benchmark. We instantiate the protocol in WirelessMathBench-XL, a 4,027-problem wireless mathematical-reasoning benchmark built from 836 retained arXiv papers across 20 subfields. Against 12.9B streamed 13-grams from RedPajama-arXiv, the audit identifies a strict zero-hit view S0 covering 3,853 problems (95.7%). Filtering to S0 changes accuracy by less than 1 pp for every evaluated model; frontier calibration rows form one high-accuracy cluster between 86.5% and 91.3%, not a resolved rank order. Only 30/800 test items carry detected overlap. Under an all-flagged-correct counterfactual, their largest possible positive score inflation is 0.31-0.51 pp for the frontier rows, so full-versus-S0 is a bounded, structurally underpowered stability summary rather than a contamination-effect test or cleanliness claim. The audit channel does not cover paraphrase, target-answer, post-training, or closed-corpus exposure. The release includes source-paper identifiers, verifier-facing ground truths, audit and threshold metadata, filtered views, a paper-disjoint sensitivity view, evaluation traces, paired-bootstrap scripts, training recipes, Croissant metadata, and a Datasheet for Datasets.
Figures & tables
Figure 1: WirelessMathBench-XL construction pipeline. We filter wireless-relevant arXiv papers, extract structured mathematical records from LaTeX sources, generate three verifier-facing task formats, and apply automated and expert QA before attaching audit metadata per problem. Counts inside the MCQ, Fill-in, and FEC boxes denote pre-QA generated candidates; the final release contains 4,027 accepted problems.
n
S0 count
S0 %
Flip vs. 13
Post-snap. %
Pre-snap. %
8
264/800
33.0
63.2
64.5
74.9
10
600/800
75.0
21.2
23.6
29.2
12
745/800
93.1
3.1
6.8
7.2
13
770/800
96.2
0.0
3.1
5.6
15
790/800
98.8
2.5
1.3
1.0
Table 1: Sensitivity of the strict zero-hit label to n -gram length on the 800 -item evaluation test split, with a retrospective collision-rate proxy. Post-snap. is the flag rate among test problems whose source papers postdate the RedPajama-ArXiv snapshot (7 March 2023) and so cannot match their own source; Pre-snap. is the flag rate among earlier papers. The release uses n=13 as a fixed operational threshold, not a validated optimum.
Model
Size
Tok
Full %
S0 %
Δ (pp) [ 95% CI]
Public %
GPT- 5 -mini
–
16 k
91.25
91.69
−0.44[−1.05,+0.08]
90.32
GPT- 5.2 -chat
–
16 k
90.75
90.91
−0.16[−0.67,+0.26]
90.00
GPT- 4.1 -mini
–
16 k
89.62
89.35
+0.27[−0.03,+0.52]
89.35
Gemini- 2.5 -Pro
–
16 k
86.88
87.01
−0.14[−0.70,+0.33]
85.48
Gemini- 2.5 -Flash
–
16 k
86.50
86.49
+0.01[−0.49,+0.43]
84.19
Table 2: Frontier calibration at a 16 k answer budget. Full is the 800 -item test split, S0 is the 770 -item strict zero-hit subset, and Δ=Full−S0 with paired-bootstrap 95% CIs; filtering moves every row by at most 0.44 pp. Public is accuracy on the 310 -item public test subset (Appendix L ), the externally reproducible view.
Row
Model
Size
Tok
Full %
S0 %
Δ (pp) [ 95% CI]
Public %
General-purpose calibration models
D1
Grok- 4 -Fast
–
02 k
88.00
88.31
−0.31[−0.90,+0.20]
84.84
D2
DeepSeek-V 3.1
671B
02 k
84.88
84.42
+0.46[+0.14,+0.75]
83.87
D3
Claude- 4.0 -Sonnet
–
02 k
83.00
82.73
+0.27[−0.20,+0.67]
82.26
D4
o 4 -mini
–
02 k
81.88
82.08
−0.20[−0.81,+0.34]
79.35
D5
DeepSeek-R 1†
671B
02 k
79.12
79.09
+0.03[−0.55,+0.55]
77.74
Table 3: Locked 2 k calibration rows under the same full-vs- S0 protocol. Rows D1–D9 are general-purpose calibration rows and D10–D14 are open/reference rows. † marks construction-involved models retained only as descriptive calibration; they are excluded from comparative claims. Public is accuracy on the 310 -item public test subset (Appendix L ). Per-format breakdowns are in Appendix K .
n
∣S0(n)∣
S0(n) %
GPT- 5 -mini
GPT- 5.2 -chat
GPT- 4.1 -mini
Gemini-Pro
Gemini-Flash
8
264/800
33.0
−0.04
+2.49
+1.75
+1.27
+1.27
10
600/800
75.0
−0.42
−0.75
+0.96
−0.62
+0.17
12
745/800
93.1
−0.29
−0.26
+0.50
−0.10
+0.19
13
770/800
96.2
−0.44
−0.16
+0.27
−0.14
+0.01
15
790/800
98.8
−0.14
−0.14
+0.00
+0.04
+0.04
Table 4: Threshold sensitivity of the frontier full-vs-cleaned comparison. Each cell is Δ=Full−S0(n) in percentage points using released 16 k frontier verdicts; paired-bootstrap intervals are released in the accompanying CSV.
Model
MCQ
Fill-in
FEC
Overall
WirelessMathLM- 7 B + GRPO
48.12
50.63
40.84
47.88
DeepSeekMath- 7 B-RL
52.63
41.39
31.94
41.00
Qwen 2.5 -Math- 7 B-Inst.
47.37
38.03
39.79
40.00
WirelessMathLM- 3 B + GRPO
29.32
28.15
16.23
25.50
Table 5: Per-format breakdown for the locked 2 k reference rows. The judged open-form columns are descriptive release-split measurements, not paper-disjoint model rankings.
Model
PD %
Rest %
Δ (pp) [ 95% CI]
GPT- 5 -mini
97.1
87.9
+9.2[+1.9,+13.8]
Grok- 4 -Fast
82.4
88.3
−5.9[−19.2,+6.4]
DeepSeek-V 3.1
85.3
84.9
+0.4[−12.5,+11.6]
WirelessMathLM- 7 B + GRPO
58.8
47.4
+11.4[−5.6,+28.4]
DeepSeekMath- 7 B-RL
58.8
40.2
+18.6[+1.5,+35.2]
Qwen 2.5 -Math- 7 B-Instruct
38.2
40.1
−1.8[−18.3,+15.5]
Table 6: Source-paper-disjoint sensitivity on the locked 2 k evaluation table. Only 34 test items come from papers absent from the training pool, so the slice is a sensitivity view rather than a replacement leaderboard.
Base
Base %
+GRPO %
Δ (rel.)
Qwen 2.5 - 0.5 B
13.38
14.87
+1.49 ( +11% )
Qwen 2.5 - 3 B
12.37
25.12
+12.75 ( +103% )
Qwen 3 - 4 B
13.63
17.87
+4.24 ( +31% )
Qwen 2.5 - 7 B
21.88
39.50
+17.62 ( +81% )
SFT ablation, Qwen 2.5 - 3 B (peak across 6 epochs)
SFT (Epoch 1)
—
18.62
—
Table 7: Training-time GRPO checks under greedy model selection. These values are separate from the locked T=0.6 +judge calibration scores; the 0.5 B row remains near the floor and is included for completeness.
Appendix figures & tables12 assets
Supplementary material from the paper’s appendix.
Appendix
Score
Criteria
1 - Invalid
Problem statement or solution is clearly wrong or contradictory; Not related to wireless communications domain; Cannot be used as a valid question
2 - Poor
Statement correct but problem too trivial (answerable instantly); Problem too vague or nearly impossible to answer correctly; Very little learning or evaluation value
3 - Acceptable
Statement and solution reasonable with no major errors; Difficulty and relevance are average; Can be kept but adds limited value (baseline quality)
4 - Good
Clear and well-structured problem; Relevant to domain and moderately challenging; Provides meaningful assessment of understanding; Worth keeping and recommending
5 - Excellent
Highly relevant to the domain; Strong depth, creativity, or insight required; Excellent for differentiating levels of understanding; Strongly recommended for inclusion
Background: Federated fine-tuning system with low-rank adaptation matrices Ak∈Rd×r , Bk∈Rr×d Question: Which term completes: W+[MASK] ? Options: A) AkBk B) BkAk Human: 4/5
LLM Score: 4/5 Strengths: Clear structure, complete context, accurate answer, rigorous notation Weaknesses: Could benefit from brief explanation of low-rank adaptation significance Agreement: ✓
Medium Quality
Background: H-NOMA system with variable definitions partially provided Question: Fill in [MASK] for the equation Human: 2/5
LLM Score: 3/5 Strengths: Wireless relevance, accurate answer Weaknesses: Ambiguous [MASK] usage, lacks clarity in instructions Technical Issues: Missing variable definitions Disagreement: LLM too optimistic
Low Quality
Background: Transformer model context with incomplete variable definitions Question: What replaces the “full key matrix”? Human: 1/5
LLM Score: 3/5 Strengths: Clear structure, accurate answer Weaknesses: Limited wireless relevance, focuses more on tensor parallelism Technical Issues: “Full key matrix” not defined Bias: LLM shows optimistic scoring pattern
Appendix
Table 9: Illustrative LLM-assisted annotation checks across quality levels.
Model Evaluated
LLM Cases
GPT-4.1-mini
Claude-3-Haiku
Gemini-2.0-Flash
GPT-5 base
270
41.11%
56.67%
42.22%
Gemini-2.5 Pro
183
53.01%
55.74%
56.28%
Claude-Sonnet-4
293
41.64%
60.75%
45.05%
Appendix
Table 10: Cross-evaluator agreement on the LLM-judged subset. GPT-5 (base) is included as a verifier-bias probe target, not as a current leaderboard row (see Tabs. 2 and 3 ).
#
Model
Size
MCQ
Fill-in
FEC
Overall
1
GPT- 5.2 -chat
–
93.98
89.08
92.67
90.75
2
GPT- 5 -mini
–
92.48
86.97
88.48
88.25
3
Grok- 4 -Fast
–
87.22
88.24
87.96
88.00
4
DeepSeek-V 3.1
671B
85.71
85.71
82.20
84.88
5
Claude-Sonnet- 4
–
84.21
84.03
79.58
83.00
6
o 4 -mini
–
83.46
83.82
75.92
81.88
Appendix
Table 11: Locked 2 k per-format evaluation on the 800 -item release test split. Rows are ranked by overall accuracy; bold marks the best score in each format. Gemini- 2.5 -Pro † is retained as a locked-budget calibration row, but Appendix J reports 2 k truncation on 3.4% of its items.
Model i
ΔS0 (pp)
95% CI (pp)
GPT- 5 -mini
+39.74
[+35.97,+43.51]
Grok- 4 -Fast
+39.61
[+35.71,+43.51]
DeepSeek-V 3.1
+35.71
[+31.56,+39.87]
Claude- 4.0 -Sonnet
+34.03
[+30.00,+37.92]
o 4 -mini
+33.38
[+29.48,+37.40]
Gemini- 2.5 -Flash
+31.82
[+27.79,+35.84]
Appendix
Table 12: Selected paired-bootstrap comparisons against WirelessMathLM- 7 B on the strict zero-hit subset. The comparison uses 10,000 resamples over the 770 paired S0 test items under the locked 2 k protocol.
WirelessMathLM
MMLU
GPQA
IFEval
HumanEval
Qwen 2.5 - 3 B-Base
42.49
23.74
26.99
61.59
+ GRPO
45.30
28.28
32.90
65.24
Δ
+2.81
+4.54
+5.91
+3.65
Qwen 2.5 - 7 B-Base
38.84
28.79
36.06
63.41
+ GRPO
50.74
38.89
35.67
64.63
Δ
+11.90
+10.10
−0.39
+1.22
Appendix
Table 13: General-purpose benchmark check after WirelessMathBench-XL GRPO. The only negative delta is −0.39 pp on IFEval for the 7 B model.
WirelessMathLM
MATH- 500
Minerva-Math
OlympiadBench
AMC
AIME- 24
Avg.
Qwen 2.5 - 7 B-Base
52.00
12.13
25.33
27.71
6.67
24.77
+ GRPO
67.00
14.34
30.22
40.96
13.33
33.17
Δ
+15.00
+2.21
+4.89
+13.25
+6.66
+8.40
Qwen 2.5 - 3 B-Base
41.60
5.88
14.67
18.07
0.00
16.04
+ GRPO
58.20
9.93
22.96
21.69
0.00
22.56
Δ
+16.60
+4.05
+8.29
+3.62
0.00
+6.52
Appendix
Table 14: Descriptive general-math check after WirelessMathBench-XL GRPO. No general-math problems appear in WirelessMathLM training data, but the Math12K GRPO control was not scored on these five benchmarks; the rows are therefore regression checks rather than an attribution claim.
ID
Model
Tok
Full %
Public %
95% CI
Δ (pp)
–
GPT- 5 -mini
16 k
91.25
90.32
[86.77,93.55]
−0.93
–
GPT- 5.2 -chat
16 k
90.75
90.00
[86.45,93.23]
−0.75
–
GPT- 4.1 -mini
16 k
89.62
89.35
[85.81,92.58]
−0.27
–
Gemini- 2.5 -Pro
16 k
86.88
85.48
[81.61,89.35]
−1.39
–
Gemini- 2.5 -Flash
16 k
86.50
84.19
[80.00,88.06]
−2.31
D1
Grok- 4 -Fast
02 k
88.00
84.84
[80.65,88.71]
−3.16
Appendix
Table 15: Accuracy on the full 800 -item test split and on the 310 -item public test subset, re-scored from archived per-item verdicts. CI is a 95% bootstrap interval over public-subset items ( 10,000 resamples, seed 0 ); Δ=Public−Full . Frontier rows (top) use a 16 k answer budget; D1–D14 are the locked- 2 k rows of Tab. 3 . † marks construction-involved models.
Model
ntotal
nvalid
Empty-body %
Acc. on nvalid
gemini-3.1-pro-preview
800
604
24.5%
89.6%
gemini-3-flash-preview
800
450
43.8%
89.3%
Appendix
Table 16: Gemini-3 preview rows excluded from the headline calibration tables. Both preview APIs return many empty response bodies; conditional accuracy is computed only over non-empty responses.
Category
Count
%
Base (7B)
+GRPO
Δ Abs
Δ Rel
Channel Estimation
160
20.0%
22.50%
36.25%
+13.75%
+61.1%
Beamforming
113
14.1%
18.58%
43.36%
+24.78%
+133.4%
Resource Allocation
110
13.8%
21.82%
39.09%
+17.27%
+79.2%
MIMO/Massive MIMO
104
13.0%
23.08%
43.27%
+20.19%
+87.5%
Information Theory
103
12.9%
21.36%
46.60%
+25.24%
+118.2%
RIS/IRS
98
12.3%
17.35%
33.67%
+16.32%
+94.1%
Appendix
Table 17: Problem-level performance by multi-label subdomain tag on the 800-problem test set. This is a descriptive release-split breakdown, not a paper-disjoint or memorisation analysis.
Hyperparameter
Value
Base Model
Qwen2.5-Base {0.5B, 3B, 7B}
+ Qwen3-4B †
Optimizer
AdamW
Learning Rate
1×10−6
Weight Decay
0.1
LR Scheduler
Cosine Annealing
Appendix
Table 18: Hyperparameters for GRPO Training
Threat
Current treatment
Remaining limitation / next step
Prompt-surface exact overlap
Reverse-probe RedPajama-arXiv audit with fixed 13 -gram threshold; release ships S0/S1 labels and rebuild scripts.
The bound is channel-specific and threshold-specific; users can tighten or replace the released filter.
Answer-surface exposure
correct_answer is deliberately excluded from the v 1.0 prompt-surface audit; a narrow answer-side 13 -gram lower-bound check is reported separately.
A fuller target-equation / answer-surface audit is needed before claiming anything about answer exposure.
Paraphrase or near-duplicate exposure
Limitations are stated in the main text and audit appendix.
Approximate near-duplicate, prompt-rewrite, or embedding-style probes are complementary future audit channels, not v 1.0 results.
Closed corpora and alternate arXiv mirrors
RedPajama-arXiv is the release-defining public slice; OpenWebMath is off-domain corroboration and algebraic-stack is a negative control.
Closed/proprietary mixtures and unprobed mirrors can contain exposure that public-corpus labels miss.
Verifier / judge dependence
Exact matching and canonicalisation precede a bounded GPT-4.1-mini semantic fallback; cross-judge variation and an exact-format lower-bound check are reported.
The 4.6 – 7.3% variation is not a human/CAS precision-recall estimate; held-out verifier calibration remains future work.
Release-split dependence
All records retain paper_id ; a 34 -item source-paper-disjoint sensitivity view is materialized in the release.
v 1.0 training rows are release-split learnability checks, not source-paper-disjoint generalisation estimates.
Appendix
Table 19: Threats to validity and how they are handled in the v 1.0 release.
We present a new approach for benchmarking Large Language Model (LLM) capabilities on research-level mathematics. Existing benchmarks largely rely on static, hand-curated sets of contest or textbook-style problems as proxies for mathematical research. Instead, we establish an updatable benchmark evaluating models directly on the latest research results in mathematics. This consists of an automatic pipeline that extracts lemmas from arXiv and rewrites them into self-contained statements by making all assumptions and required definitions explicit. It results in a benchmark that can be updated regularly with new problems taken directly from human mathematical research, while previous instances can be used for training without compromising future evaluations. We benchmark current state-of-the-art LLMs, which obtain around 10-15% accuracy in theorem proving (pass@1) depending on the model, showing that there is currently a large margin of progression for LLMs to reach human-level proving capabilities in a research context.
Antoine Peyronnet, Fabian Gloeckle, Amaury Hayat
1Ecole Normale Superieure de Rennes · 2Ecole des Ponts Paris · 3Korean Institute for Advanced Study
Mathematical reasoning is a hallmark of human intelligence, and whether large language models (LLMs) can meaningfully perform it remains a central question in artificial intelligence and cognitive science. As LLMs are increasingly integrated into scientific workflows, rigorous evaluation of their mathematical capabilities becomes a practical necessity. Existing benchmarks are limited by synthetic settings and data contamination. We present LiveMathematicianBench, a dynamic multiple-choice benchmark for research-level mathematical reasoning built from recent arXiv papers published after model training cutoffs. By grounding evaluation in newly published theorems, it provides a realistic testbed beyond memorized patterns. The benchmark introduces a thirteen-category logical taxonomy of theorem types (e.g., implication, equivalence, existence, uniqueness), enabling fine-grained evaluation across reasoning forms. It employs a proof-sketch-guided distractor pipeline that uses high-level proof strategies to construct plausible but invalid answer choices reflecting misleading proof directions, increasing sensitivity to genuine understanding over surface-level matching. We also introduce a substitution-resistant mechanism to distinguish answer recognition from substantive reasoning. Evaluation shows the benchmark is far from saturated: Gemini-3.1-pro-preview, the best model, achieves only 43.5%. Under substitution-resistant evaluation, accuracy drops sharply: GPT-5.4 scores highest at 30.6%, while Gemini-3.1-pro-preview falls to 17.6%, below the 20% random baseline. A dual-mode protocol reveals that proof-sketch access yields consistent accuracy gains, suggesting models can leverage high-level proof strategies for reasoning. Overall, LiveMathematicianBench offers a scalable, contamination-resistant testbed for studying research-level mathematical reasoning in LLMs.
Linyang He, Qiyao Yu, Hanze Dong +5
Columbia University · Microsoft Research · University of Amsterdam
Large language models (LLMs) are becoming increasingly capable mathematical collaborators, but static benchmarks are no longer sufficient for evaluating progress: they are often narrow in scope, quickly saturated, and rarely updated. This makes it hard to compare models reliably and track progress over time. Instead, we need evaluation platforms: continuously maintained systems that run, aggregate, and analyze evaluations across many benchmarks to give a comprehensive picture of model performance within a broad domain. In this work, we build on the original MathArena benchmark by substantially broadening its scope from final-answer olympiad problems to a continuously maintained evaluation platform for mathematical reasoning with LLMs. MathArena now covers a much wider range of tasks, including proof-based competitions, research-level arXiv problems, and formal proof generation in Lean. Additionally, we maintain a clear evaluation protocol for all models and regularly design new benchmarks as model capabilities improve to ensure that MathArena remains challenging. Notably, the strongest model, GPT-5.5, now reaches 98% on the 2026 USA Math Olympiad and 74% on research-level questions, showing that frontier models can now comfortably solve extremely challenging mathematical problems. This highlights the importance of continuously maintained evaluation platforms like MathArena to track the rapid progress of LLMs in mathematical reasoning.
Jasper Dekoninck, Nikola Jovanović, Tim Gehrunger +4
ETH Zurich · INSAIT · Sofia University "St. Kliment Ohridski"