The Endless Exam: Mathematical Constructions from Today's Models toward Superintelligence
Organizations: The University of Texas at Arlington
Abstract
We introduce the Endless Exam, a benchmark spanning fourteen parameterised families of mathematical construction problems, with verifiable scores that distinguish progress before and beyond published mathematical frontiers. Each submitted object is checked automatically for validity and assigned a relative quality score against a published frontier or construction baseline, without capping improvements at 1. The benchmark draws long-term challenges from open mathematical problems and generates larger instances by varying their parameters. Compact certificates allow large constructions to be verified without listing every element. Across nine models evaluated on 69 distinct instances, continuous quality scores distinguish performance even though none of the 30 published-frontier references is surpassed. Size-quality curves show how construction quality changes as problem size increases. We release the generators, verifiers, references, model responses and analysis to support continued measurement before and beyond human frontiers.
Figures & tables
| Dataset | Problems | Difficulty | Unsolved | Non-proof | Open eval. | Auto- check | Uncapped score |
|---|---|---|---|---|---|---|---|
| FrontierMath | 338 | Univ.–Res. | |||||
| FrontierMath: Open Problems | 50 | Research | |||||
| IMProofBench | 77 | Research | ∗ | ||||
| First Proof | 10 | Research | |||||
| Erdős Problems | 1,217 | Research | |||||
| IMO-AnswerBench | 400 | Olympiad |
| Mean relative quality | |||||
| configuration | Overall score | Published frontiers (30) | Construction baselines (39) | Valid fraction | Gap closed |
| Qwen3.5-4B high | 7.55 | 0.07 [0.03, 0.11] | 0.08 [0.03, 0.15] | 0.38 | 0.000 |
| Qwen3.5-9B high | 14.52 | 0.12 [0.07, 0.17] | 0.17 [0.07, 0.27] | 0.43 | 0.000 |
| Qwen3.5-27B high | 16.18 | 0.13 [0.06, 0.22] | 0.18 [0.09, 0.29] | 0.52 | 0.000 |
| Qwen3.8-27B high | 38.07 | 0.31 [0.20, 0.42] | 0.44 [0.28, 0.59] | 0.62 | 0.004 |
| GPT-5.6 Luna medium | 16.70 | 0.16 [0.10, 0.24] | 0.17 [0.08, 0.27] | 0.57 | 0.000 |
Appendix figures & tables22 assets
Supplementary material from the paper’s appendix.
Appendix
| family | construction | objective | class | reference; bound or target |
| cap set | ; no three collinear points (cap set) | R | published; proven bound | |
| spherical codes | vectors with exact coordinates; pairwise angle ; or variable | count | R | mixed; proven bound |
| matrix multiplication | bilinear scheme for by matrices; computes the product exactly | rank (min) | S | published; proven bound |
| degree–diameter | graph with max degree ; diameter | vertices | S | published; proven bound |
| covering | -subsets (blocks) of ; every -subset lies in a block | blocks (min) | S | published; proven bound |
| Schur | -colouring of ; no monochromatic | R | published; proven bound |
| Family | Parameters | Reference | Baseline |
|---|---|---|---|
| cap set | d=6 | 112 | 112 |
| cap set | d=8 | 512 | 512 |
| cap set | d=10 | 2432 | 2240 |
| cap set | d=12 | 12928 | 12544 |
| spherical codes ( ) | d=8 | 240 | 240 |
| spherical codes ( ) | d=10 | 510 | 510 |
| Family | Size | Qwen3.8-27B high | GPT-5.6 Luna high | DeepSeek V4.1 Flash low | GPT-6 Astra high |
|---|---|---|---|---|---|
| cap set | 0.00 (0/2) | 0.36 (1/2) | 0.65 (2/2) | 1.00 (2/2) | |
| cap set | 0.32 (1/2) | 0.78 (2/2) | 0.57 (2/2) | 0.91 (2/2) | |
| cap set | 0.63 (2/2) | 0.63 (2/2) | 0.60 (2/2) | 0.88 (2/2) | |
| cap set | 0.31 (1/2) | 0.51 (2/2) | 0.41 (2/2) | 0.97 (2/2) | |
| spherical codes ( ) | 1.00 (2/2) | 0.00 (0/2) | 1.00 (2/2) | 1.00 (2/2) | |
| spherical codes ( ) | 0.30 (1/2) | 0.67 (2/2) | 0.82 (2/2) | 0.98 (2/2) |
| configuration | Shannon relative quality | valid | Trifference relative quality | valid | words at |
|---|---|---|---|---|---|
| Qwen3.5-4B high | 0.053 | 4/4 | 0.000 | 0/4 | 0 |
| Qwen3.5-9B high | 0.053 | 4/4 | 0.000 | 0/4 | 0 |
| Qwen3.5-27B high | 0.297 | 4/4 | 0.019 | 2/4 | 0 |
| Qwen3.8-27B high | 0.504 | 3/4 | 0.501 | 3/4 | |
| GPT-5.6 Luna medium | 0.044 | 3/4 | 0.000 | 0/4 | 0 |
| GPT-5.6 Luna high | 0.705 | 4/4 | 0.250 | 1/4 |
| Model / effort | ||||
|---|---|---|---|---|
| Qwen3.5-4B high | 0.145 | 0.037 | 0.023 | 0.007 |
| Qwen3.5-9B high | 0.145 | 0.037 | 0.023 | 0.007 |
| Qwen3.5-27B high | 0.145 | 0.037 | 1.000 | 0.007 |
| Qwen3.8-27B high | 0.713 | 0.304 | 0.000 † | 1.000 |
| GPT-5.6 Luna medium | 0.145 | 0.000 † | 0.023 | 0.007 |
| GPT-5.6 Luna high | 0.515 | 0.304 | 1.000 | 1.000 |
| Model / effort | 64 | 96 | 144 | 192 |
|---|---|---|---|---|
| Qwen3.5-4B high | 0.000 † | 0.000 † | 0.000 † | 0.000 † |
| Qwen3.5-9B high | 0.000 † | 0.000 † | 0.000 † | 0.000 † |
| Qwen3.5-27B high | 0.037 | 0.037 | 0.000 † | 0.000 † |
| Qwen3.8-27B high | 1.000 | 1.000 | 0.000 † | 0.004 |
| GPT-5.6 Luna medium | 0.000 † | 0.000 † | 0.000 † | 0.000 † |
| GPT-5.6 Luna high | 0.000 † | 0.000 † | 0.000 † | 1.000 |
| Length | Separation | Reference | Model | Relative quality | Valid |
|---|---|---|---|---|---|
| 30 | 1 | 729 | 729 | 1.000 | yes |
| 36 | 2 | 729 | 729 | 1.000 | yes |
| 42 | 3 | 729 | 729 | 1.000 | yes |
| 56 | 1 | 2187 | 6561 | 3.000 | yes |
| Mean relative quality | ||||
|---|---|---|---|---|
| configuration | Overall score | Published frontiers (15) | Construction baselines (30) | Valid fraction |
| Claude Opus 5 medium | 52.41 | 0.48 | 0.55 | 0.87 |
| Qwen3.5-4B high | 7.69 | 0.08 | 0.07 | 0.38 |
| Qwen3.5-9B high | 16.96 | 0.14 | 0.18 | 0.47 |
| Qwen3.5-27B high | 17.07 | 0.17 | 0.17 | 0.53 |
| Qwen3.8-27B high | 40.98 | 0.37 | 0.43 | 0.64 |
| configuration | Relative quality | Relative quality | Gap closed | Gap closed |
|---|---|---|---|---|
| Qwen3.5-4B high | 0.088 | 0.096 | 0.000 | 0.000 |
| Qwen3.5-9B high | 0.191 | 0.202 | 0.000 | 0.000 |
| Qwen3.5-27B high | 0.184 | 0.192 | 0.000 | 0.000 |
| Qwen3.8-27B high | 0.425 | 0.418 | 0.005 | 0.004 |
| GPT-5.6 Luna medium | 0.203 | 0.207 | 0.000 | 0.000 |
| GPT-5.6 Luna high | 0.419 | 0.408 | 0.002 | 0.002 |
| Gap closed to proven bounds | LABS target | |||
| Configuration | All 64 | Published 30 | Construction 34 | 5 instances |
| Qwen3.5-4B high | 0.000 | 0.000 | 0.000 | 0.000 |
| Qwen3.5-9B high | 0.000 | 0.000 | 0.000 | 0.000 |
| Qwen3.5-27B high | 0.000 | 0.000 | 0.000 | 0.000 |
| Qwen3.8-27B high | 0.004 | 0.000 | 0.007 | 0.000 |
| GPT-5.6 Luna medium | 0.000 | 0.000 | 0.000 | 0.000 |
| test construction | words | primary (s) | independent (s) |
|---|---|---|---|
| rank 12, | 531,441 | 5.61 | 8.66 |
| rank 12, | 531,441 | 21.63 | 37.70 |
| concatenation, |
| Reference group | Verified instances | Largest answer (tokens) |
|---|---|---|
| Published frontiers | 30 / 30 | 77,528 |
| Construction baselines | 39 / 39 | 19,149 |
| Total | 69 / 69 | 77,528 |
| Family / parameters | Reference and source | Bound and source |
|---|---|---|
| Cap set, | 2432, 5504; Edel and Karapetyan–Karapetyan constructions ( Edel, 2004 ; Karapetyan & Karapetyan, 2023 ) . | 5619, 16857; proven bounds from Versluis (2017) and slicing. |
| Kissing, | 1154, 1932, 2564; constructions ( Zinoviev & Ericson, 1999 ; Ganzhinov, 2025 ; Leech & Sloane, 1971 ) . | 2064, 3174, 4853; proven bounds ( Leijenhorst & de Laat, 2024 ) . |
| Matrix multiplication, , | 61, 93; schemes from the Lille FMM catalogue. | 20, 25; proven lower bounds from tensor flattenings. |
| Degree–diameter, , , , | 364, 168, 196, 104; graphs from Comellas (2026) . | 485, 302, 382, 161; proven Moore bounds ( Miller & Širáň, 2013 ) . |
| AP-free, | 194, 649; affine charts of Edel’s projective caps ( Edel, 2026 ; Elsholtz & Pach, 2020 ) . | 1296, 7497; proven bounds from the polynomial-method count ( Ellenberg & Gijswijt, 2017 ) . |
| MOLS, | 2, 4, 5; constructions ( Miller et al., 2024 ; Todorov, 2012 ; Abel, 2015 ) . | 8, 12, 17; proven or bounds, including Lam et al. (1989) . |
| Family / parameters | Reference and source | Bound or target and source |
|---|---|---|
| Kakeya control, , , | Expert constructions of sizes 703, 553, 2741 ( Saraf & Sudan, 2008 ) ; construction baselines. | 595, 375, 2105; proven lower bounds ( Bukh & Chao, 2021 ) . |
| LABS | Rotated Legendre construction or reference search ( Packebusch & Mertens, 2016 ) . | Merit factor 12.32; conjectured target. |
| Heilbronn (both domains) | Expert construction or reference search; see Appendix A . | ; trivial proven bound from triangulation. |
| Linear-equation-free sets; corners | Digit constructions or reference search ( Behrend, 1946 ; Ruzsa, 1993 ) . | and ; trivial proven bounds. |
| Spherical codes | Expert construction or reference search. | Cap-packing bound; proven. |
| instance | length | reference words | verified words | Relative quality | generation (s) |
|---|---|---|---|---|---|
| T1 | 64 | 19,683 | 19,683 | 1.000 | 52.3 |
| T2 | 96 | 531,441 | 531,441 | 1.000 | 53.8 |
| T3 | 144 | 387,420,489 | 387,420,489 | 1.000 | 37.0 |
| T4 | 192 | 10,460,353,203 | 10,460,353,203 | 1.000 | 63.8 |
| family | class | runs | instances | mean | max | mean | max |
| cap set | R | 3 | 2 | 0.108 | 0.163 | 0.000 | 0.000 |
| spherical codes | R | 6 | 5 | 0.000 | 0.000 | 0.000 | 0.000 |
| corners | R | 3 | 3 | 0.019 | 0.024 | 0.000 | 0.000 |
| linear-equation-free sets | R | 3 | 3 | 0.016 | 0.019 | 0.016 | 0.019 |
| Schur | R | 3 | 3 | 0.000 | 0.000 | 0.000 | 0.000 |
| AP-free | R | 3 | 2 | 0.041 | 0.083 | 0.041 | 0.083 |
| Family | Parameters |
|---|---|
| Cap sets | |
| Cap sets | |
| Cap sets | |
| Corners | |
| Corners | |
| Corners |
| Domain | Vertex 1 | Vertex 2 | Vertex 3 |
|---|---|---|---|
| Family | Tier | Representative parameters |
|---|---|---|
| Cap sets | A1 | ; ; |
| A2 | ; ; | |
| Corner-free sets | A1 | ; ; |
| A2 | ; ; | |
| Schur colourings | A1 | ; ; |
| A2 | ; ; |
| Astra high | Luna high | Opus 5.5 high | |||||
|---|---|---|---|---|---|---|---|
| Family | No tools | Tools | No tools | Tools | No tools | Tools | |
| cap set | 3 | 0.923 | 0.984 | 0.464 | 0.946 | 0.877 | 0.955 |
| LABS | 5 | 0.985 | 1.374 | 0.000 | 1.189 | 0.379 | 1.166 |
| Heilbronn | 10 | 0.684 | 2.138 | 0.196 | 2.270 | 0.456 | 2.936 |
| spherical codes | 8 | 1.313 | 1.642 | 0.284 | 1.420 | 0.927 | 1.784 |
| corners | 5 | 1.026 | 1.109 | 0.961 | 1.078 | 1.009 | 1.117 |
| Measure | Astra high | Luna high | Opus 5.5 high |
|---|---|---|---|
| Input tokens | 34,575,392 | 98,019,732 | 157,011,011 ∗ |
| Cached input tokens (included above) | 31,453,056 | 91,690,496 | 149,741,729 ∗ |
| Output tokens | 445,123 | 866,731 | 1,702,589 ∗ |
| Reasoning tokens (included above) | 181,416 | 425,435 | 668,352 ∗ |
| Astra high | Luna high | Opus 5.5 high | ||||||
|---|---|---|---|---|---|---|---|---|
| Family | Parameters | Ref. | No tools | Tools | No tools | Tools | No tools | Tools |
| cap set | P | 0.932 | 1.000 | 0.739 | 1.000 | 0.932 | 0.946 | |
| cap set | P | 0.921 | 1.000 | 0.000 † | 0.921 | 0.785 | 0.921 | |
| cap set | P | 0.916 | 0.951 | 0.654 | 0.916 | 0.916 | 0.997 | |
| LABS | C | 0.931 | 1.568 | 0.000 † | 1.306 | 0.931 | 1.067 | |
| LABS | C | 0.985 | 1.507 | 0.000 † | 1.253 | 0.000 † | 1.123 | |