cs.CLMar 3, 2026

Code2Math: Can Your Code Agent Evolve Math Problems Through Exploration?

Authors: Dadi Guo, Yuejin Xie, Qingyu Liu, Weixian Huang, Jiayu Liu, Zhiyuan Fan, Qihan Ren, Shuai Shao, +6 more

Organizations: Hong Kong University of Science and Technology · Tsinghua University · Zhejiang University · Nanjing Tech University · Shanghai Jiao Tong University · University of Michigan · Independent Researcher

Abstract

As large language models (LLMs) advance their mathematical capabilities toward the IMO and research level, the scarcity of challenging, high-quality problems has become a significant bottleneck for training, evaluation and self-evolution of LLMs. Simultaneously, recent code agents have demonstrated sophisticated skills in agentic coding and reasoning, suggesting that code execution can serve as a scalable environment for mathematical experimentation. In this paper, we investigate the potential of code agents to autonomously evolve existing math problems into more complex variations. We introduce a multi-agent framework designed to perform problem evolution while validating the solvability and increased difficulty of the generated problems. Our experiments demonstrate that, given sufficient test-time exploration, code agents can synthesize new, solvable problems that are structurally distinct from and more challenging than the originals. This work provides empirical evidence that code-driven agents can serve as a viable mechanism for synthesizing high-difficulty mathematical reasoning problems within scalable computational environments. Code and data is available at https://github.com/TarferSoul/Code2Math.

Figures & tables

Appendix figures & tables1 asset

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

May 27, 2026cs.CL

ResearchMath-14K: Scaling Research-Level Mathematics via Agents

The frontier of mathematics is defined by problems whose solutions are not yet known. However, whether language models can meaningfully engage with such problems without human intervention remains unclear. A major obstacle is the lack of large-scale research-level math datasets. To this end, we introduce ResearchMath-14k, a set of 14,05614{,}056 problems curated from academic sources via a multi-agent pipeline. ResearchMath-14k spans 11 mathematical domains and ranks above existing math datasets on knowledge, novelty, and procedural difficulty. To our knowledge, it is the largest research-level mathematical problem set available for training. We additionally generate 220220K teacher trajectories through targeted prompting, followed by behavioral filtering. Notably, however, generating correct trajectories is nontrivial at this level, and two LLM judges label only 3.7%3.7\% and 4.3%4.3\% of sampled ResearchMath training trajectories as correct. Nevertheless, across three model families, full-parameter training on ResearchMath improves performance on graduate- and research-level mathematics benchmarks by 2.12.1 points over the starting checkpoints. In comparison, training on existing datasets such as DASD and Nemotron-SFT-Math-v4 changes performance by 0.00.0 and −0.5-0.5 points, respectively. Notably, mixing DASD with ResearchMath yields higher scores than token-matched DASD alone on benchmarks covering olympiad short-form (+2.0+2.0), graduate- and research-level short-form (+0.8+0.8), graduate- and research-level symbolic (+2.6+2.6), and proof evaluation (+5.7+5.7). Further analysis suggests that research-level mathematical content and greater reasoning diversity may help explain why ResearchMath provides complementary supervision to contemporary datasets. We make ResearchMath-14k publicly available for future works on research-level mathematical reasoning.
May 26, 2026cs.CL

ReverseMath: Answer Inversion for Scalable and Verifiable Mathematical Problem Generation

Mathematical reasoning benchmarks are vital for evaluating large language models (LLMs), but many are static and repeatedly exposed through public evaluation and training pipelines, making it difficult to separate genuine reasoning from memorization. Meanwhile, manually constructing new math problems with reliable answers remains costly. We introduce ReverseMath, a scalable method for generating new math problems through answer inversion. Given a problem and its answer, ReverseMath masks a numerical value in the original problem, treats the original answer as a known condition, and rewrites the problem so that the masked value becomes the new answer. The generated problem reverses the original input-output relation, making its answer known by construction. We study ReverseMath for both evaluation and training. For evaluation, paired original/reversed problems reveal substantial behavioral shifts: models sometimes fail on reversed problems and even incorrectly output the original answer, suggesting memorization-like behavior. For training, ReverseMath provides automatically labeled reversed problems as data augmentation for reinforcement learning (RL). Experiments show that including ReverseMath-generated data improves mathematical reasoning performance across multiple benchmarks, demonstrating its value as both an analysis tool and a scalable source of verifiable training data.
Jul 10, 2026cs.AI

ProofCouncil: An LLM Agent for Solving Open Mathematical Problems

Large language models (LLMs) have shown increasing promise in solving open problems in mathematics. However, their performance can be further improved through agentic workflows tailored to real-world mathematical practice. To this end, we introduce ProofCouncil, a mathematical agent that is designed to tackle open problems using an author-critic architecture. ProofCouncil served as a submission to the second batch of FirstProof, a challenge consisting of 10 real-world mathematical problems that agents must solve autonomously. Its submissions for 6 of the 10 problems were judged by the referees to be correct up to at most minor revisions, showing the best performance among participating teams. We also evaluate ProofCouncil on 30 open problems collected from mathematical researchers. Among the 21 solutions that received human feedback, 5 were judged completely correct, 2 more were judged promising pending final verification, and a further 8 contained useful partial progress. In this short paper, we describe the development of ProofCouncil and the agent-building library used to create it, which we release as open source to the community.