Organizations: Department of Modern Physics, University of Science and Technology of China, Hefei, Anhui 230026, China · CAS Key Laboratory of Theoretical Physics, Institute of Theoretical Physics, Chinese Academy of Sciences, Beijing 100190, China · Hefei National Laboratory, University of Science and Technology of China, Hefei 230088, China · Hefei National Research Center for Physical Sciences at the Microscale and School of Physical Sciences, University of Science and Technology of China, Hefei 230026, China · Institute for Advanced Algorithms Research, Shanghai 200120, China · Endless Frontier, Shanghai 200030, China
The next frontier for artificial general intelligence is tackling unresolved scientific problems, calling for benchmarks that assess progress beyond established knowledge. We introduce OpenProblemBench, a benchmark of 82 unresolved problems drawn from the mathematics and theoretical physics literature. Each problem supplies the research context, assumptions, and prior progress needed to investigate the question. We select problems whose proposed solutions admit comparatively clear checks of their decisive mathematical or computational claims. Four evaluator models independently assess the correctness, completeness, and degree of progress of each submission without reference solutions. Across seven evaluated configurations, GPT-6-Astra achieves the highest mean judged solve rate of 14.0%, compared with 5.5-6.7% for the evaluated full-size open models and 2.4-3.7% for Flash models. Case comparisons connect stronger outcomes to changes in problem representation, general arguments that extend beyond finite evidence, and proofs of the steps needed to complete a solution. By grounding evaluation in questions arising from the research literature, OpenProblemBench provides a setting for investigating the capabilities and limitations of AI as a contributor to foundational theoretical science.
Figures & tables
Figure 1: Outcome distribution averaged equally across the four evaluators and ordered by mean judge-assessed solve rate. Each evaluator uses all 82 problems as the denominator. The final segment combines completed nonsolved verdicts with incomplete assessments, for which the evaluator reported insufficient evidence to make a resolution judgment. Raw counts retain these two outcomes separately.
Figure 2: Solve rates under each evaluator for the seven primary configurations. Each cell uses 82 problems; incomplete assessments contribute no successes. Execution environments are specified in Section 3.3 .
Figure 3: Partial progress and overclaim diagnostics under each evaluator. Left: nontrivial plus breakthrough outcomes among completed reviews not judged solved. Right: detected overclaims among reviews with a determined overclaim verdict; diagnostic coverage is reported in Appendix B .
Harness
Web access
Solved n (%)
Nontrivial + breakthrough n (%)
OpenCode
No
7 (8.54)
68 (82.93)
Claude Code
No
5 (6.10)
59 (71.95)
Claude Code
Yes
5 (6.10)
68 (82.93)
Table 1: Qwen3.8-Max harness and web-access comparisons on the same 82 statements, with GLM-5.3 using the same offline review protocol. Web access refers to the solver. Nontrivial + breakthrough counts partial results only; all percentages use 82 as the denominator.
Appendix figures & tables9 assets
Supplementary material from the paper’s appendix.
Appendix
Benchmark
Evaluation target
Evidence / assessment
SciCode; ScienceAgentBench
Research coding and publication-derived workflows
Reference implementations, tests, and execution outputs
FrontierScience
Expert scientific tasks and research subtasks
Scientist-authored answers and granular criteria
AstaBench; AutoResearchEval; TB-Science
Scientific workflows and agent behavior
Task-specific outcomes, tools, and process analysis
FIRE-Bench; TruthInsightBench
Insight rediscovery / evidence-grounded discovery on study-derived tasks
Replication of machine-learning papers / exploratory physics tasks
Author-developed rubrics / reference answers and expert rubrics
MLRC-Bench
Open-ended machine-learning methodology
Empirical research performance
Appendix
Table 2: Representative evaluation designs, distinguishing tasks with published targets from unresolved research questions.
Model
GPT
Kimi
GLM
Qwen
GPT-6-Astra
15 (18.29%)
13 (15.85%)
9 (10.98%)
9 (10.98%)
GPT-5.6-Sol
7 (8.54%)
6 (7.32%)
5 (6.10%)
5 (6.10%)
Kimi-K3
3 (3.66%)
7 (8.54%)
5 (6.10%)
5 (6.10%)
GLM-5.3
1 (1.22%)
6 (7.32%)
6 (7.32%)
5 (6.10%)
Qwen3.8-Max
2 (2.44%)
7 (8.54%)
7 (8.54%)
6 (7.32%)
DeepSeek-V4.1-Flash
1 (1.22%)
2 (2.44%)
3 (3.66%)
2 (2.44%)
Appendix
Table 3: Submissions judged solved out of 82, with percentages in parentheses. All cells contain 82 review records; incomplete outcome assessments are reported separately.
Model
GPT
Kimi
GLM
Qwen
GPT-6-Astra
60/2/67
48/1/69
62/2/73
54/2/72
GPT-5.6-Sol
49/1/75
30/2/76
54/2/77
43/3/77
Kimi-K3
62/9/78
39/3/74
57/4/77
45/4/76
GLM-5.3
57/1/80
38/4/74
49/3/76
34/2/71
Qwen3.8-Max
56/2/79
41/4/75
62/6/75
46/2/70
DeepSeek-V4.1-Flash
17/2/65
19/1/79
30/0/79
16/0/74
Appendix
Table 4: Nontrivial / breakthrough / all partial counts. Categories include completed assessments only; these are review events, not unique discoveries.
Model
GPT
Kimi
GLM
Qwen
GPT-6-Astra
7/63/12
0/82/0
2/80/0
4/78/0
GPT-5.6-Sol
7/69/6
3/79/0
6/75/1
4/76/2
Kimi-K3
38/44/0
33/49/0
60/22/0
54/28/0
GLM-5.3
8/17/57
13/69/0
59/23/0
57/23/2
Qwen3.8-Max
16/14/52
16/66/0
61/21/0
52/25/5
DeepSeek-V4.1-Flash
18/0/64
30/52/0
67/15/0
47/35/0
Appendix
Table 5: Overclaim counts: detected / not detected / unknown. Unknown means that the archived diagnostic did not establish a complete yes/no verdict, usually because coverage was partial. Labels are recovered from linked diagnostic files or explicit archived report verdicts. A dash means no reviews for the entire cell.
Model
GPT
Kimi
GLM
Qwen
GPT-6-Astra
85.4%
100.0%
100.0%
100.0%
GPT-5.6-Sol
92.7%
100.0%
98.8%
97.6%
Kimi-K3
100.0%
100.0%
100.0%
100.0%
GLM-5.3
30.5%
100.0%
100.0%
97.6%
Qwen3.8-Max
36.6%
100.0%
100.0%
93.9%
DeepSeek-V4.1-Flash
22.0%
100.0%
100.0%
100.0%
Appendix
Table 6: Overclaim diagnostic coverage: determined verdicts as a percentage of 82 submissions. The complement is the unknown share.
Figure 4: Numbers of problems judged solved by zero, one to three, or four evaluators for each configuration. The zero category includes submissions with Incomplete assessments and no solved judgment.
Problem ID
Astra
GPT-5.6
Kimi
GLM
Qwen
GLM Flash
DeepSeek Flash
ORB-PHYS-57
S
N
S
B
N
N
T
ORB-PHYS-62
S
N
N
N
N
N
T
ORB-PHYS-69
S
S
B
N
B
N
N
ORB-PHYS-82
S
S
S
S
S
S
S
Appendix
Table 7: Contrasting approaches on four research problems, assessed by the same Qwen reviewer. S: solved; B: breakthrough partial; N: nontrivial partial; T: trivial partial. Model abbreviations refer to the primary configurations in Section 3.3. Problem content and proof obligations are described below.
Mathematics
n
Physics
n
Combinatorics and discrete mathematics
12
Statistical mechanics and integrable models
21
Geometry and topology
8
Algebraic/combinatorial mathematical physics
7
Algebraic geometry and commutative algebra
7
Quantum information and computation
4
Analysis, partial differential equations, and dynamics
6
Field theory, gauge theory, and strings
3
Number theory and arithmetic dynamics
5
Statistical optics
1
Representation theory and operator algebras
4
Appendix
Table 8: Primary-subfield composition. Each problem is assigned to one primary subfield for counting; the accompanying data give all assignments.
ID
Problem
ORB-MATH-01
Irrationality and arithmetic nature of the Euler–Mascheroni constant
ORB-MATH-02
Griffiths’ 1969 positive-polynomial problem at the level of Chern forms: pointwise (weak) positivity of Schur forms for Griffiths-positive vector bundles
ORB-MATH-03
Algebraicity of Weil classes on polarized abelian varieties of Weil type
ORB-MATH-04
Quillen’s conjecture Qn : are all projective modules over the localization of a regular local ring at a regular parameter free?
ORB-MATH-05
Binary Goldbach conjecture: is every even integer greater than 2 a sum of two primes?
ORB-MATH-06
The Gorenstein bimodule conjecture: is a finite-dimensional algebra selfinjective exactly when its regular bimodule is Gorenstein projective?
Appendix
Table 9: The 82 retained problem statements. Titles are inherited from the archive. Domain prefixes record collection provenance.
Can AI make progress on important, unsolved mathematical problems? Large language models are now capable of sophisticated mathematical and scientific reasoning, but whether they can perform novel research is still widely debated and underexplored. We introduce HorizonMath, a benchmark of 113 predominantly unsolved problems spanning eight domains in mathematics and the mathematical sciences, paired with an open-source evaluation framework for automated verification. Our benchmark targets the generator-verifier gap: problems where discovery is hard and requires meaningful mathematical insight, but verification is computationally straightforward. This contrasts with most existing research-level benchmarks, which instead rely on formal proof verification or manual review, both of which are expensive to scale. Because these solutions are unknown, HorizonMath is resistant to data contamination, and most state-of-the-art models score under 10%. Using this framework, we identify six novel solutions to research problems that either resolve previously open questions or improve on the best-known published results, with GPT-5.4 Pro and GPT-5.6 Sol each discovering three of these solutions. Across seven frontier model families, reasoning efficiency and behavior also vary substantially. We release HorizonMath as an open challenge and a growing community resource, where each verified solution is a candidate contribution to the mathematical literature.
Erik Y. Wang, Sumeet R. Motwani, James V. Roggeveen +9
Stanford University · Benchmark · University of Oxford +3
Recent AI systems have achieved gold-medal-level performance on the International Mathematical Olympiad, demonstrating remarkable proficiency at competition-style problem solving. However, competition mathematics represents only a narrow slice of mathematical reasoning: problems are drawn from limited domains, require minimal advanced machinery, and can often reward insightful tricks over deep theoretical knowledge. We introduce Riemann-Bench, a private benchmark of expert-curated problems designed to evaluate AI systems on research-level mathematics that goes far beyond the olympiad frontier. Problems are authored by Ivy League mathematics professors, graduate students, and PhD-holding IMO medalists, and routinely took their authors weeks to solve independently. Each problem undergoes double-blind verification by two independent domain experts who must solve the problem from scratch, and yields a unique, closed-form solution assessed by programmatic verifiers. We evaluate frontier models as unconstrained research agents, with full access to coding tools, search, and open-ended reasoning, using an unbiased statistical estimator computed over 100 independent runs per problem. Our results reveal that all frontier models currently score below 10%, exposing a substantial gap between olympiad-level problem solving and genuine research-level mathematical reasoning. By keeping the benchmark fully private, we ensure that measured performance reflects authentic mathematical capability rather than memorization of training data.
Following the recent achievement of gold-medal performance on the IMO by frontier LLMs, the community is searching for the next meaningful and challenging target for measuring LLM reasoning. Whereas olympiad-style problems measure step-by-step reasoning alone, research-level problems use such reasoning to advance the frontier of mathematical knowledge itself, emerging as a compelling alternative. Yet research-level math benchmarks remain scarce because such problems are difficult to source (e.g., Riemann Bench and FrontierMath-Tier 4 contain 25 and 50 problems, respectively). To support reliable evaluation of next-generation frontier models, we introduce Soohak, a 439-problem benchmark newly authored from scratch by 64 mathematicians. Soohak comprises two subsets. On the Challenge subset, frontier models including Gemini-3-Pro, GPT-5, and Claude-Opus-4.5 reach 30.4%, 26.4%, and 10.4% respectively, leaving substantial headroom, while leading open-weight models such as Qwen3-235B, GPT-OSS-120B, and Kimi-2.5 remain below 15%. Notably, beyond standard problem solving, Soohak introduces a refusal subset that probes a capability intrinsic to research mathematics: recognizing ill-posed problems and pausing rather than producing confident but unjustified answers. On this subset, no model exceeds 50%, identifying refusal as a new optimization target that current models do not directly address. To prevent contamination, the dataset will be publicly released in late 2026, with model evaluations available upon request in the interim.