WirelessMathBench-XL: An Auditable Benchmark for Wireless Mathematical Reasoning
Organizations: Nanyang Technological University
Abstract
Technical-domain benchmarks constructed from arXiv papers can overlap the same public text used in LLM pretraining. Auditing this risk at training-corpus scale requires searching billions of corpus n-grams while retaining per-item evidence that users can inspect and recompute. We contribute a reverse-probe audit at a fixed 13-gram threshold: it indexes benchmark prompts, streams public pretraining corpora, and emits per-problem prompt-surface lexical-overlap metadata with memory that scales with the benchmark. We instantiate the protocol in WirelessMathBench-XL, a 4,027-problem wireless mathematical-reasoning benchmark built from 836 retained arXiv papers across 20 subfields. Against 12.9B streamed 13-grams from RedPajama-arXiv, the audit identifies a strict zero-hit view S0 covering 3,853 problems (95.7%). Filtering to S0 changes accuracy by less than 1 pp for every evaluated model; frontier calibration rows form one high-accuracy cluster between 86.5% and 91.3%, not a resolved rank order. Only 30/800 test items carry detected overlap. Under an all-flagged-correct counterfactual, their largest possible positive score inflation is 0.31-0.51 pp for the frontier rows, so full-versus-S0 is a bounded, structurally underpowered stability summary rather than a contamination-effect test or cleanliness claim. The audit channel does not cover paraphrase, target-answer, post-training, or closed-corpus exposure. The release includes source-paper identifiers, verifier-facing ground truths, audit and threshold metadata, filtered views, a paper-disjoint sensitivity view, evaluation traces, paired-bootstrap scripts, training recipes, Croissant metadata, and a Datasheet for Datasets.
Figures & tables
| count | % | Flip vs. | Post-snap. % | Pre-snap. % | |
|---|---|---|---|---|---|
| Model | Size | Tok | Full % | % | (pp) [ CI] | Public % |
|---|---|---|---|---|---|---|
| GPT- -mini | – | k | 91.25 | 91.69 | 90.32 | |
| GPT- -chat | – | k | 90.75 | 90.91 | 90.00 | |
| GPT- -mini | – | k | 89.62 | 89.35 | 89.35 | |
| Gemini- -Pro | – | k | 86.88 | 87.01 | 85.48 | |
| Gemini- -Flash | – | k | 86.50 | 86.49 | 84.19 |
| Row | Model | Size | Tok | Full % | % | (pp) [ CI] | Public % |
|---|---|---|---|---|---|---|---|
| General-purpose calibration models | |||||||
| D1 | Grok- -Fast | – | k | 88.00 | 88.31 | 84.84 | |
| D2 | DeepSeek-V | 671B | k | 84.88 | 84.42 | 83.87 | |
| D3 | Claude- -Sonnet | – | k | 83.00 | 82.73 | 82.26 | |
| D4 | o -mini | – | k | 81.88 | 82.08 | 79.35 | |
| D5 | DeepSeek-R | 671B | k | 79.12 | 79.09 | 77.74 | |
| % | GPT- -mini | GPT- -chat | GPT- -mini | Gemini-Pro | Gemini-Flash | ||
|---|---|---|---|---|---|---|---|
| 8 | 264/800 | 33.0 | |||||
| 10 | 600/800 | 75.0 | |||||
| 12 | 745/800 | 93.1 | |||||
| 13 | 770/800 | 96.2 | |||||
| 15 | 790/800 | 98.8 |
| Model | MCQ | Fill-in | FEC | Overall |
|---|---|---|---|---|
| WirelessMathLM- B + GRPO | 48.12 | 50.63 | 40.84 | 47.88 |
| DeepSeekMath- B-RL | 52.63 | 41.39 | 31.94 | 41.00 |
| Qwen -Math- B-Inst. | 47.37 | 38.03 | 39.79 | 40.00 |
| WirelessMathLM- B + GRPO | 29.32 | 28.15 | 16.23 | 25.50 |
| Model | PD % | Rest % | (pp) [ CI] |
|---|---|---|---|
| GPT- -mini | 97.1 | 87.9 | |
| Grok- -Fast | 82.4 | 88.3 | |
| DeepSeek-V | 85.3 | 84.9 | |
| WirelessMathLM- B + GRPO | 58.8 | 47.4 | |
| DeepSeekMath- B-RL | 58.8 | 40.2 | |
| Qwen -Math- B-Instruct | 38.2 | 40.1 |
| Base | Base % | +GRPO % | (rel.) |
|---|---|---|---|
| Qwen - B | 13.38 | 14.87 | +1.49 ( ) |
| Qwen - B | 12.37 | 25.12 | +12.75 ( ) |
| Qwen - B | 13.63 | 17.87 | +4.24 ( ) |
| Qwen - B | 21.88 | 39.50 | +17.62 ( ) |
| SFT ablation, Qwen - B (peak across epochs) | |||
| SFT (Epoch 1) | — | 18.62 | — |
Appendix figures & tables12 assets
Supplementary material from the paper’s appendix.
Appendix
| Score | Criteria |
|---|---|
| 1 - Invalid | Problem statement or solution is clearly wrong or contradictory; Not related to wireless communications domain; Cannot be used as a valid question |
| 2 - Poor | Statement correct but problem too trivial (answerable instantly); Problem too vague or nearly impossible to answer correctly; Very little learning or evaluation value |
| 3 - Acceptable | Statement and solution reasonable with no major errors; Difficulty and relevance are average; Can be kept but adds limited value (baseline quality) |
| 4 - Good | Clear and well-structured problem; Relevant to domain and moderately challenging; Provides meaningful assessment of understanding; Worth keeping and recommending |
| 5 - Excellent | Highly relevant to the domain; Strong depth, creativity, or insight required; Excellent for differentiating levels of understanding; Strongly recommended for inclusion |
| Quality Level | Question Content | LLM Assessment |
|---|---|---|
| High Quality | Background: Federated fine-tuning system with low-rank adaptation matrices , Question: Which term completes: ? Options: A) B) Human: 4/5 | LLM Score: 4/5 Strengths: Clear structure, complete context, accurate answer, rigorous notation Weaknesses: Could benefit from brief explanation of low-rank adaptation significance Agreement: |
| Medium Quality | Background: H-NOMA system with variable definitions partially provided Question: Fill in [MASK] for the equation Human: 2/5 | LLM Score: 3/5 Strengths: Wireless relevance, accurate answer Weaknesses: Ambiguous [MASK] usage, lacks clarity in instructions Technical Issues: Missing variable definitions Disagreement: LLM too optimistic |
| Low Quality | Background: Transformer model context with incomplete variable definitions Question: What replaces the “full key matrix”? Human: 1/5 | LLM Score: 3/5 Strengths: Clear structure, accurate answer Weaknesses: Limited wireless relevance, focuses more on tensor parallelism Technical Issues: “Full key matrix” not defined Bias: LLM shows optimistic scoring pattern |
| Model Evaluated | LLM Cases | GPT-4.1-mini | Claude-3-Haiku | Gemini-2.0-Flash |
|---|---|---|---|---|
| GPT-5 base | 270 | 41.11% | 56.67% | 42.22% |
| Gemini-2.5 Pro | 183 | 53.01% | 55.74% | 56.28% |
| Claude-Sonnet-4 | 293 | 41.64% | 60.75% | 45.05% |
| # | Model | Size | MCQ | Fill-in | FEC | Overall |
|---|---|---|---|---|---|---|
| 1 | GPT- -chat | – | 93.98 | 89.08 | 92.67 | 90.75 |
| 2 | GPT- -mini | – | 92.48 | 86.97 | 88.48 | 88.25 |
| 3 | Grok- -Fast | – | 87.22 | 88.24 | 87.96 | 88.00 |
| 4 | DeepSeek-V | 671B | 85.71 | 85.71 | 82.20 | 84.88 |
| 5 | Claude-Sonnet- | – | 84.21 | 84.03 | 79.58 | 83.00 |
| 6 | o -mini | – | 83.46 | 83.82 | 75.92 | 81.88 |
| Model | (pp) | CI (pp) |
|---|---|---|
| GPT- -mini | ||
| Grok- -Fast | ||
| DeepSeek-V | ||
| Claude- -Sonnet | ||
| o -mini | ||
| Gemini- -Flash |
| WirelessMathLM | MMLU | GPQA | IFEval | HumanEval |
|---|---|---|---|---|
| Qwen - B-Base | 42.49 | 23.74 | 26.99 | 61.59 |
| + GRPO | 45.30 | 28.28 | 32.90 | 65.24 |
| +2.81 | +4.54 | +5.91 | +3.65 | |
| Qwen - B-Base | 38.84 | 28.79 | 36.06 | 63.41 |
| + GRPO | 50.74 | 38.89 | 35.67 | 64.63 |
| +11.90 | +10.10 | +1.22 |
| WirelessMathLM | MATH- | Minerva-Math | OlympiadBench | AMC | AIME- | Avg. |
|---|---|---|---|---|---|---|
| Qwen - B-Base | 52.00 | 12.13 | 25.33 | 27.71 | 6.67 | 24.77 |
| + GRPO | 67.00 | 14.34 | 30.22 | 40.96 | 13.33 | 33.17 |
| +15.00 | +2.21 | +4.89 | +13.25 | +6.66 | +8.40 | |
| Qwen - B-Base | 41.60 | 5.88 | 14.67 | 18.07 | 0.00 | 16.04 |
| + GRPO | 58.20 | 9.93 | 22.96 | 21.69 | 0.00 | 22.56 |
| +16.60 | +4.05 | +8.29 | +3.62 | 0.00 | +6.52 |
| ID | Model | Tok | Full % | Public % | CI | (pp) |
|---|---|---|---|---|---|---|
| – | GPT- -mini | k | 91.25 | 90.32 | ||
| – | GPT- -chat | k | 90.75 | 90.00 | ||
| – | GPT- -mini | k | 89.62 | 89.35 | ||
| – | Gemini- -Pro | k | 86.88 | 85.48 | ||
| – | Gemini- -Flash | k | 86.50 | 84.19 | ||
| D1 | Grok- -Fast | k | 88.00 | 84.84 |
| Model | Empty-body % | Acc. on | ||
|---|---|---|---|---|
| gemini-3.1-pro-preview | 800 | 604 | ||
| gemini-3-flash-preview | 800 | 450 |
| Category | Count | % | Base (7B) | +GRPO | Abs | Rel |
|---|---|---|---|---|---|---|
| Channel Estimation | 160 | 20.0% | 22.50% | 36.25% | +13.75% | +61.1% |
| Beamforming | 113 | 14.1% | 18.58% | 43.36% | +24.78% | +133.4% |
| Resource Allocation | 110 | 13.8% | 21.82% | 39.09% | +17.27% | +79.2% |
| MIMO/Massive MIMO | 104 | 13.0% | 23.08% | 43.27% | +20.19% | +87.5% |
| Information Theory | 103 | 12.9% | 21.36% | 46.60% | +25.24% | +118.2% |
| RIS/IRS | 98 | 12.3% | 17.35% | 33.67% | +16.32% | +94.1% |
| Hyperparameter | Value |
|---|---|
| Base Model | Qwen2.5-Base {0.5B, 3B, 7B} |
| + Qwen3-4B † | |
| Optimizer | AdamW |
| Learning Rate | |
| Weight Decay | 0.1 |
| LR Scheduler | Cosine Annealing |
| Threat | Current treatment | Remaining limitation / next step |
|---|---|---|
| Prompt-surface exact overlap | Reverse-probe RedPajama-arXiv audit with fixed -gram threshold; release ships labels and rebuild scripts. | The bound is channel-specific and threshold-specific; users can tighten or replace the released filter. |
| Answer-surface exposure | correct_answer is deliberately excluded from the v prompt-surface audit; a narrow answer-side -gram lower-bound check is reported separately. | A fuller target-equation / answer-surface audit is needed before claiming anything about answer exposure. |
| Paraphrase or near-duplicate exposure | Limitations are stated in the main text and audit appendix. | Approximate near-duplicate, prompt-rewrite, or embedding-style probes are complementary future audit channels, not v results. |
| Closed corpora and alternate arXiv mirrors | RedPajama-arXiv is the release-defining public slice; OpenWebMath is off-domain corroboration and algebraic-stack is a negative control. | Closed/proprietary mixtures and unprobed mirrors can contain exposure that public-corpus labels miss. |
| Verifier / judge dependence | Exact matching and canonicalisation precede a bounded GPT-4.1-mini semantic fallback; cross-judge variation and an exact-format lower-bound check are reported. | The – variation is not a human/CAS precision-recall estimate; held-out verifier calibration remains future work. |
| Release-split dependence | All records retain paper_id ; a -item source-paper-disjoint sensitivity view is materialized in the release. | v training rows are release-split learnability checks, not source-paper-disjoint generalisation estimates. |