Organizations: Department of Modern Physics, University of Science and Technology of China, Hefei, Anhui 230026, China · CAS Key Laboratory of Theoretical Physics, Institute of Theoretical Physics, Chinese Academy of Sciences, Beijing 100190, China · Hefei National Laboratory, University of Science and Technology of China, Hefei 230088, China · Hefei National Research Center for Physical Sciences at the Microscale and School of Physical Sciences, University of Science and Technology of China, Hefei 230026, China · Institute for Advanced Algorithms Research, Shanghai 200120, China · Endless Frontier, Shanghai 200030, China
The next frontier for artificial general intelligence is tackling unresolved scientific problems, calling for benchmarks that assess progress beyond established knowledge. We introduce OpenProblemBench, a benchmark of 82 unresolved problems drawn from the mathematics and theoretical physics literature. Each problem supplies the research context, assumptions, and prior progress needed to investigate the question. We select problems whose proposed solutions admit comparatively clear checks of their decisive mathematical or computational claims. Four evaluator models independently assess the correctness, completeness, and degree of progress of each submission without reference solutions. Across seven evaluated configurations, GPT-6-Astra achieves the highest mean judged solve rate of 14.0%, compared with 5.5-6.7% for the evaluated full-size open models and 2.4-3.7% for Flash models. Case comparisons connect stronger outcomes to changes in problem representation, general arguments that extend beyond finite evidence, and proofs of the steps needed to complete a solution. By grounding evaluation in questions arising from the research literature, OpenProblemBench provides a setting for investigating the capabilities and limitations of AI as a contributor to foundational theoretical science.
Figures & tables
Figure 1: Outcome distribution averaged equally across the four evaluators and ordered by mean judge-assessed solve rate. Each evaluator uses all 82 problems as the denominator. The final segment combines completed nonsolved verdicts with incomplete assessments, for which the evaluator reported insufficient evidence to make a resolution judgment. Raw counts retain these two outcomes separately.
Figure 2: Solve rates under each evaluator for the seven primary configurations. Each cell uses 82 problems; incomplete assessments contribute no successes. Execution environments are specified in Section 3.3 .
Figure 3: Partial progress and overclaim diagnostics under each evaluator. Left: nontrivial plus breakthrough outcomes among completed reviews not judged solved. Right: detected overclaims among reviews with a determined overclaim verdict; diagnostic coverage is reported in Appendix B .
Harness
Web access
Solved n (%)
Nontrivial + breakthrough n (%)
OpenCode
No
7 (8.54)
68 (82.93)
Claude Code
No
5 (6.10)
59 (71.95)
Claude Code
Yes
5 (6.10)
68 (82.93)
Table 1: Qwen3.8-Max harness and web-access comparisons on the same 82 statements, with GLM-5.3 using the same offline review protocol. Web access refers to the solver. Nontrivial + breakthrough counts partial results only; all percentages use 82 as the denominator.
Appendix figures & tables9 assets
Supplementary material from the paper’s appendix.
Appendix
Benchmark
Evaluation target
Evidence / assessment
SciCode; ScienceAgentBench
Research coding and publication-derived workflows
Reference implementations, tests, and execution outputs
FrontierScience
Expert scientific tasks and research subtasks
Scientist-authored answers and granular criteria
AstaBench; AutoResearchEval; TB-Science
Scientific workflows and agent behavior
Task-specific outcomes, tools, and process analysis
FIRE-Bench; TruthInsightBench
Insight rediscovery / evidence-grounded discovery on study-derived tasks
Replication of machine-learning papers / exploratory physics tasks
Author-developed rubrics / reference answers and expert rubrics
MLRC-Bench
Open-ended machine-learning methodology
Empirical research performance
Appendix
Table 2: Representative evaluation designs, distinguishing tasks with published targets from unresolved research questions.
Model
GPT
Kimi
GLM
Qwen
GPT-6-Astra
15 (18.29%)
13 (15.85%)
9 (10.98%)
9 (10.98%)
GPT-5.6-Sol
7 (8.54%)
6 (7.32%)
5 (6.10%)
5 (6.10%)
Kimi-K3
3 (3.66%)
7 (8.54%)
5 (6.10%)
5 (6.10%)
GLM-5.3
1 (1.22%)
6 (7.32%)
6 (7.32%)
5 (6.10%)
Qwen3.8-Max
2 (2.44%)
7 (8.54%)
7 (8.54%)
6 (7.32%)
DeepSeek-V4.1-Flash
1 (1.22%)
2 (2.44%)
3 (3.66%)
2 (2.44%)
Appendix
Table 3: Submissions judged solved out of 82, with percentages in parentheses. All cells contain 82 review records; incomplete outcome assessments are reported separately.
Model
GPT
Kimi
GLM
Qwen
GPT-6-Astra
60/2/67
48/1/69
62/2/73
54/2/72
GPT-5.6-Sol
49/1/75
30/2/76
54/2/77
43/3/77
Kimi-K3
62/9/78
39/3/74
57/4/77
45/4/76
GLM-5.3
57/1/80
38/4/74
49/3/76
34/2/71
Qwen3.8-Max
56/2/79
41/4/75
62/6/75
46/2/70
DeepSeek-V4.1-Flash
17/2/65
19/1/79
30/0/79
16/0/74
Appendix
Table 4: Nontrivial / breakthrough / all partial counts. Categories include completed assessments only; these are review events, not unique discoveries.
Model
GPT
Kimi
GLM
Qwen
GPT-6-Astra
7/63/12
0/82/0
2/80/0
4/78/0
GPT-5.6-Sol
7/69/6
3/79/0
6/75/1
4/76/2
Kimi-K3
38/44/0
33/49/0
60/22/0
54/28/0
GLM-5.3
8/17/57
13/69/0
59/23/0
57/23/2
Qwen3.8-Max
16/14/52
16/66/0
61/21/0
52/25/5
DeepSeek-V4.1-Flash
18/0/64
30/52/0
67/15/0
47/35/0
Appendix
Table 5: Overclaim counts: detected / not detected / unknown. Unknown means that the archived diagnostic did not establish a complete yes/no verdict, usually because coverage was partial. Labels are recovered from linked diagnostic files or explicit archived report verdicts. A dash means no reviews for the entire cell.
Model
GPT
Kimi
GLM
Qwen
GPT-6-Astra
85.4%
100.0%
100.0%
100.0%
GPT-5.6-Sol
92.7%
100.0%
98.8%
97.6%
Kimi-K3
100.0%
100.0%
100.0%
100.0%
GLM-5.3
30.5%
100.0%
100.0%
97.6%
Qwen3.8-Max
36.6%
100.0%
100.0%
93.9%
DeepSeek-V4.1-Flash
22.0%
100.0%
100.0%
100.0%
Appendix
Table 6: Overclaim diagnostic coverage: determined verdicts as a percentage of 82 submissions. The complement is the unknown share.
Figure 4: Numbers of problems judged solved by zero, one to three, or four evaluators for each configuration. The zero category includes submissions with Incomplete assessments and no solved judgment.
Problem ID
Astra
GPT-5.6
Kimi
GLM
Qwen
GLM Flash
DeepSeek Flash
ORB-PHYS-57
S
N
S
B
N
N
T
ORB-PHYS-62
S
N
N
N
N
N
T
ORB-PHYS-69
S
S
B
N
B
N
N
ORB-PHYS-82
S
S
S
S
S
S
S
Appendix
Table 7: Contrasting approaches on four research problems, assessed by the same Qwen reviewer. S: solved; B: breakthrough partial; N: nontrivial partial; T: trivial partial. Model abbreviations refer to the primary configurations in Section 3.3. Problem content and proof obligations are described below.
Mathematics
n
Physics
n
Combinatorics and discrete mathematics
12
Statistical mechanics and integrable models
21
Geometry and topology
8
Algebraic/combinatorial mathematical physics
7
Algebraic geometry and commutative algebra
7
Quantum information and computation
4
Analysis, partial differential equations, and dynamics
6
Field theory, gauge theory, and strings
3
Number theory and arithmetic dynamics
5
Statistical optics
1
Representation theory and operator algebras
4
Appendix
Table 8: Primary-subfield composition. Each problem is assigned to one primary subfield for counting; the accompanying data give all assignments.
ID
Problem
ORB-MATH-01
Irrationality and arithmetic nature of the Euler–Mascheroni constant
ORB-MATH-02
Griffiths’ 1969 positive-polynomial problem at the level of Chern forms: pointwise (weak) positivity of Schur forms for Griffiths-positive vector bundles
ORB-MATH-03
Algebraicity of Weil classes on polarized abelian varieties of Weil type
ORB-MATH-04
Quillen’s conjecture Qn : are all projective modules over the localization of a regular local ring at a regular parameter free?
ORB-MATH-05
Binary Goldbach conjecture: is every even integer greater than 2 a sum of two primes?
ORB-MATH-06
The Gorenstein bimodule conjecture: is a finite-dimensional algebra selfinjective exactly when its regular bimodule is Gorenstein projective?
Appendix
Table 9: The 82 retained problem statements. Titles are inherited from the archive. Domain prefixes record collection provenance.