cs.AISep 21, 2026

The Endless Exam: Mathematical Constructions from Today's Models toward Superintelligence

Authors: Muhan Zhang

Organizations: The University of Texas at Arlington

Abstract

We introduce the Endless Exam, a benchmark spanning fourteen parameterised families of mathematical construction problems, with verifiable scores that distinguish progress before and beyond published mathematical frontiers. Each submitted object is checked automatically for validity and assigned a relative quality score against a published frontier or construction baseline, without capping improvements at 1. The benchmark draws long-term challenges from open mathematical problems and generates larger instances by varying their parameters. Compact certificates allow large constructions to be verified without listing every element. Across nine models evaluated on 69 distinct instances, continuous quality scores distinguish performance even though none of the 30 published-frontier references is surpassed. Size-quality curves show how construction quality changes as problem size increases. We release the generators, verifiers, references, model responses and analysis to support continued measurement before and beyond human frontiers.

Figures & tables

Appendix figures & tables22 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. ResearchMath-14K: Scaling Research-Level Mathematics via Agents

    May 27, 2026Guijin Son, Seungyeop Yi, Minju Gwak +3Research-Level MathematicsMathematics

  2. MathConstraint: Automated Generation of Verified Combinatorial Reasoning Instances for LLMs

    May 8, 2026Viresh Pati, Zhengyu Li, Piyush Jha +3CombinationsConstraint Satisfaction