cs.AIOct 8, 2026

OpenProblemBench: Benchmarking AI on Open Problems in the Foundational Theoretical Sciences

Authors: Zhiyi Li, Sihan Hu, Tianning Xiao, Xiansheng Cai, Xiaojun Tan, Youjin Deng, Kun Chen

Organizations: Department of Modern Physics, University of Science and Technology of China, Hefei, Anhui 230026, China · CAS Key Laboratory of Theoretical Physics, Institute of Theoretical Physics, Chinese Academy of Sciences, Beijing 100190, China · Hefei National Laboratory, University of Science and Technology of China, Hefei 230088, China · Hefei National Research Center for Physical Sciences at the Microscale and School of Physical Sciences, University of Science and Technology of China, Hefei 230026, China · Institute for Advanced Algorithms Research, Shanghai 200120, China · Endless Frontier, Shanghai 200030, China

Abstract

The next frontier for artificial general intelligence is tackling unresolved scientific problems, calling for benchmarks that assess progress beyond established knowledge. We introduce OpenProblemBench, a benchmark of 82 unresolved problems drawn from the mathematics and theoretical physics literature. Each problem supplies the research context, assumptions, and prior progress needed to investigate the question. We select problems whose proposed solutions admit comparatively clear checks of their decisive mathematical or computational claims. Four evaluator models independently assess the correctness, completeness, and degree of progress of each submission without reference solutions. Across seven evaluated configurations, GPT-6-Astra achieves the highest mean judged solve rate of 14.0%, compared with 5.5-6.7% for the evaluated full-size open models and 2.4-3.7% for Flash models. Case comparisons connect stronger outcomes to changes in problem representation, general arguments that extend beyond finite evidence, and proofs of the steps needed to complete a solution. By grounding evaluation in questions arising from the research literature, OpenProblemBench provides a setting for investigating the capabilities and limitations of AI as a contributor to foundational theoretical science.

Figures & tables

Appendix figures & tables9 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Measuring Progress in Reasoning Toward Mathematical Discovery with Automatic Verification

    Date pendingErik Y. Wang, Sumeet R. Motwani, James V. Roggeveen +9Mathematical Reasoning BenchmarksLLM Mathematical Reasoning