cs.AISep 21, 2026

The Endless Exam: Mathematical Constructions from Today's Models toward Superintelligence

Authors: Muhan Zhang

Organizations: The University of Texas at Arlington

Abstract

We introduce the Endless Exam, a benchmark spanning fourteen parameterised families of mathematical construction problems, with verifiable scores that distinguish progress before and beyond published mathematical frontiers. Each submitted object is checked automatically for validity and assigned a relative quality score against a published frontier or construction baseline, without capping improvements at 1. The benchmark draws long-term challenges from open mathematical problems and generates larger instances by varying their parameters. Compact certificates allow large constructions to be verified without listing every element. Across nine models evaluated on 69 distinct instances, continuous quality scores distinguish performance even though none of the 30 published-frontier references is surpassed. Size-quality curves show how construction quality changes as problem size increases. We release the generators, verifiers, references, model responses and analysis to support continued measurement before and beyond human frontiers.

Figures & tables

Appendix figures & tables22 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

May 27, 2026cs.CL

ResearchMath-14K: Scaling Research-Level Mathematics via Agents

The frontier of mathematics is defined by problems whose solutions are not yet known. However, whether language models can meaningfully engage with such problems without human intervention remains unclear. A major obstacle is the lack of large-scale research-level math datasets. To this end, we introduce ResearchMath-14k, a set of 14,05614{,}056 problems curated from academic sources via a multi-agent pipeline. ResearchMath-14k spans 11 mathematical domains and ranks above existing math datasets on knowledge, novelty, and procedural difficulty. To our knowledge, it is the largest research-level mathematical problem set available for training. We additionally generate 220220K teacher trajectories through targeted prompting, followed by behavioral filtering. Notably, however, generating correct trajectories is nontrivial at this level, and two LLM judges label only 3.7%3.7\% and 4.3%4.3\% of sampled ResearchMath training trajectories as correct. Nevertheless, across three model families, full-parameter training on ResearchMath improves performance on graduate- and research-level mathematics benchmarks by 2.12.1 points over the starting checkpoints. In comparison, training on existing datasets such as DASD and Nemotron-SFT-Math-v4 changes performance by 0.00.0 and −0.5-0.5 points, respectively. Notably, mixing DASD with ResearchMath yields higher scores than token-matched DASD alone on benchmarks covering olympiad short-form (+2.0+2.0), graduate- and research-level short-form (+0.8+0.8), graduate- and research-level symbolic (+2.6+2.6), and proof evaluation (+5.7+5.7). Further analysis suggests that research-level mathematical content and greater reasoning diversity may help explain why ResearchMath provides complementary supervision to contemporary datasets. We make ResearchMath-14k publicly available for future works on research-level mathematical reasoning.
May 8, 2026cs.LG

MathConstraint: Automated Generation of Verified Combinatorial Reasoning Instances for LLMs

We introduce MathConstraint, a hard, adaptive benchmark for evaluating the combinatorial reasoning capabilities of LLMs. We combine constraint satisfaction problems with rigorous solver-based verification and design an adaptive generator to create instances that remain challenging as the LLMs improve in their reasoning capabilities. Unlike existing benchmarks that quickly saturate on fixed datasets or use LLM-as-a-judge for checking solutions,MathConstraint uses parameterized problem types that enable scalable generation of arbitrarily difficult and automatically verifiable instances. We release MathConstraint-Easy (266266 instances), on which frontier models achieve between 72.6%72.6\% (gemini-3.1-flash-lite) and 87.6%87.6\% (gpt-5.5) accuracy, and MathConstraint (329329 instances) on which the same models drop to between 18.5%18.5\% (claude-4.6-sonnet) and 66.9%66.9\% (gpt-5.5) accuracy, demonstrating the resilience of our benchmark generator against rapid progress in LLM reasoning capabilities. We evaluate 12 frontier and open-weight models with and without access to a sandboxed Python environment that includes generic SAT/SMT solvers. Tool access roughly doubles frontier accuracy on MathConstraint (mean +28+28pp; up to +52+52pp for claude-4.6-sonnet). Further, halving the tool-call budget from 88 to 44 rounds erases up to 3737 points -- a sensitivity that most single-budget benchmarks miss. We release the generator, dataset, and evaluation harness as a robust environment for studying combinatorial reasoning and tool-use behavior under adversarially-tunable difficulty.
May 9, 2026cs.CL

Soohak: A Mathematician-Curated Benchmark for Evaluating Research-level Math Capabilities of LLMs

Following the recent achievement of gold-medal performance on the IMO by frontier LLMs, the community is searching for the next meaningful and challenging target for measuring LLM reasoning. Whereas olympiad-style problems measure step-by-step reasoning alone, research-level problems use such reasoning to advance the frontier of mathematical knowledge itself, emerging as a compelling alternative. Yet research-level math benchmarks remain scarce because such problems are difficult to source (e.g., Riemann Bench and FrontierMath-Tier 4 contain 25 and 50 problems, respectively). To support reliable evaluation of next-generation frontier models, we introduce Soohak, a 439-problem benchmark newly authored from scratch by 64 mathematicians. Soohak comprises two subsets. On the Challenge subset, frontier models including Gemini-3-Pro, GPT-5, and Claude-Opus-4.5 reach 30.4%, 26.4%, and 10.4% respectively, leaving substantial headroom, while leading open-weight models such as Qwen3-235B, GPT-OSS-120B, and Kimi-2.5 remain below 15%. Notably, beyond standard problem solving, Soohak introduces a refusal subset that probes a capability intrinsic to research mathematics: recognizing ill-posed problems and pausing rather than producing confident but unjustified answers. On this subset, no model exceeds 50%, identifying refusal as a new optimization target that current models do not directly address. To prevent contamination, the dataset will be publicly released in late 2026, with model evaluations available upon request in the interim.