Function-Structured Reinforcement Learning with Executable Verifiers for Mathematical Reasoning
Organizations: Department of Computer Science, University College London, UK · Institute of Software, Chinese Academy of Sciences, China
Abstract
Algorithmic mathematical reasoning requires reliable decomposition, computation, and aggregation. Final-answer rewards provide limited guidance on intermediate errors, while successful execution does not guarantee mathematical correctness. This work proposes Function-Structured Graph Reinforcement Learning (FSG-RL), connecting subproblem graphs and Python implementations with multi-verifier feedback. The policy first learns to generate code from function graphs through supervised fine-tuning (SFT). Group Relative Policy Optimization (GRPO) then optimizes the policy using answer-gated rewards and span-level credit assignment. The framework also supports teacher supervision and structured memory. A benchmark curated from Grade School Math 8K (GSM8K), MathQA, MATH, and Omni-MATH pairs public function graphs with private verification specifications. Under a unified evaluation protocol, GRPO improves final-answer accuracy from 43.25% to 67.50% and full solution success from 32.25% to 52.25% over SFT. Continued reinforcement learning (RL) with teacher supervision yields additional gains. The gains extend beyond producing correctly formatted code, supporting verifier-guided reinforcement learning for mathematical reasoning. Code is available at https://github.com/ZihanLiummyycc/FSG-RL.
Figures & tables
| Model | Exec. | Tests | Back. | Ans. | Full |
|---|---|---|---|---|---|
| Qwen Base | 23.75 | 20.00 | 1.25 | 2.50 | 1.00 |
| Executable SFT | 86.00 | 61.75 | 41.50 | 43.25 | 32.25 |
| FSG-RL | 90.25 | 65.75 | 63.75 | 67.50 | 52.25 |
| Model | Exec. | Tests | Back. | Ans. | Full |
|---|---|---|---|---|---|
| FSG-RL | 90.25 | 65.75 | 63.75 | 67.50 | 52.25 |
| FC-GRPO | 90.75 | 67.00 | 63.75 | 69.00 | 53.50 |
| Qwen3.8-27B API | 54.75 | 52.00 | 52.50 | 52.50 | 49.75 |