cs.LGSep 27, 2025

WirelessMathBench-XL: An Auditable Benchmark for Wireless Mathematical Reasoning

Authors: Xin Li, Mengbing Liu, Yiyang Zhu, Wenhe Zhang, Li Wei, Jiancheng An, Chau Yuen

Organizations: Nanyang Technological University

Abstract

Technical-domain benchmarks constructed from arXiv papers can overlap the same public text used in LLM pretraining. Auditing this risk at training-corpus scale requires searching billions of corpus n-grams while retaining per-item evidence that users can inspect and recompute. We contribute a reverse-probe audit at a fixed 13-gram threshold: it indexes benchmark prompts, streams public pretraining corpora, and emits per-problem prompt-surface lexical-overlap metadata with memory that scales with the benchmark. We instantiate the protocol in WirelessMathBench-XL, a 4,027-problem wireless mathematical-reasoning benchmark built from 836 retained arXiv papers across 20 subfields. Against 12.9B streamed 13-grams from RedPajama-arXiv, the audit identifies a strict zero-hit view S0 covering 3,853 problems (95.7%). Filtering to S0 changes accuracy by less than 1 pp for every evaluated model; frontier calibration rows form one high-accuracy cluster between 86.5% and 91.3%, not a resolved rank order. Only 30/800 test items carry detected overlap. Under an all-flagged-correct counterfactual, their largest possible positive score inflation is 0.31-0.51 pp for the frontier rows, so full-versus-S0 is a bounded, structurally underpowered stability summary rather than a contamination-effect test or cleanliness claim. The audit channel does not cover paraphrase, target-answer, post-training, or closed-corpus exposure. The release includes source-paper identifiers, verifier-facing ground truths, audit and threshold metadata, filtered views, a paper-disjoint sensitivity view, evaluation traces, paired-bootstrap scripts, training recipes, Croissant metadata, and a Datasheet for Datasets.

Figures & tables

Appendix figures & tables12 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. LemmaBench: A Live, Research-Level Benchmark to Evaluate LLM Capabilities in Mathematics

    Feb 27, 2026Antoine Peyronnet, Fabian Gloeckle, Amaury HayatResearch-Level MathematicsTheorem Proving

  2. LiveMathematicianBench: A Live Benchmark for Research-Level Mathematical Reasoning with Proof Sketches

    Apr 2, 2026Linyang He, Qiyao Yu, Hanze Dong +5Research-Level MathematicsLarge Language Model Benchmarks

  3. Beyond Benchmarks: MathArena as an Evaluation Platform for Mathematics with LLMs

    May 1, 2026Jasper Dekoninck, Nikola Jovanović, Tim Gehrunger +4Mathematical Reasoning BenchmarksMathematical Reasoning