cs.CLMar 3, 2026

Code2Math: Can Your Code Agent Evolve Math Problems Through Exploration?

Authors: Dadi Guo, Yuejin Xie, Qingyu Liu, Weixian Huang, Jiayu Liu, Zhiyuan Fan, Qihan Ren, Shuai Shao, +6 more

Organizations: Hong Kong University of Science and Technology · Tsinghua University · Zhejiang University · Nanjing Tech University · Shanghai Jiao Tong University · University of Michigan · Independent Researcher

Abstract

As large language models (LLMs) advance their mathematical capabilities toward the IMO and research level, the scarcity of challenging, high-quality problems has become a significant bottleneck for training, evaluation and self-evolution of LLMs. Simultaneously, recent code agents have demonstrated sophisticated skills in agentic coding and reasoning, suggesting that code execution can serve as a scalable environment for mathematical experimentation. In this paper, we investigate the potential of code agents to autonomously evolve existing math problems into more complex variations. We introduce a multi-agent framework designed to perform problem evolution while validating the solvability and increased difficulty of the generated problems. Our experiments demonstrate that, given sufficient test-time exploration, code agents can synthesize new, solvable problems that are structurally distinct from and more challenging than the originals. This work provides empirical evidence that code-driven agents can serve as a viable mechanism for synthesizing high-difficulty mathematical reasoning problems within scalable computational environments. Code and data is available at https://github.com/TarferSoul/Code2Math.

Figures & tables

Appendix figures & tables1 asset

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. ResearchMath-14K: Scaling Research-Level Mathematics via Agents

    May 27, 2026Guijin Son, Seungyeop Yi, Minju Gwak +3Research-Level MathematicsMathematics

  2. ReverseMath: Answer Inversion for Scalable and Verifiable Mathematical Problem Generation

    May 26, 2026Raoyuan Zhao, Yihong Liu, Yupei Du +2Mathematical Reasoning BenchmarksReverse

  3. ProofCouncil: An LLM Agent for Solving Open Mathematical Problems

    Jul 10, 2026Johannes Schmitt, Tim Gehrunger, Jasper Dekoninck +4Open ProblemsLarge Language Model Agents