Self-evolving LLM agents promise to improve autonomously through continual interaction and learning, reducing their dependence on manually curated supervision. Realizing this promise requires not only updating the agent, but also evolving its training experience as its capabilities change. However, most existing pipelines rely on static datasets or separately updated synthesis models, causing previously useful tasks to become trivial while overly difficult tasks remain uninformative. This growing mismatch between agent capability and training experience limits sustained self-improvement. To address this problem, we propose SynCo, an agentic data synthesis co-training framework for self-evolving LLMs based on multi-agent reinforcement learning. SynCo jointly optimizes two independently parameterized agents: a Synthesizer that constructs training tasks from the Reasoner's evolving capability state, and a Reasoner that learns from the resulting experience. Each synthesized task induces multiple Reasoner rollouts whose outcomes provide complementary rewards to both agents. Correctness feedback improves the Reasoner, while task quality, answer reliability, and outcome-grounded teachability guide the Synthesizer. Their updates are fed back into subsequent synthesis rounds, allowing the task-solving policy and its training distribution to evolve together. Extensive experiments across eight mathematical reasoning benchmarks demonstrate that SynCo substantially outperforms a broad range of existing synthetic-data methods and controlled baselines, achieving the strongest overall performance while deriving most of its gains from previously unsolved problems.
Figures & tables
Figure 1: Overview of SynCo , a closed-loop framework for self-evolving LLMs. (1) A Synthesizer first generates capability-guided and verifiable tasks from the current learning context. (2) The resulting interaction batch is then used to jointly update the Synthesizer and the Reasoner through online multi-agent co-training. (3) The updated agents in turn reshape the future curriculum and task distribution, forming an agent–data self-evolution loop.
Method
Primary Benchmarks
Advanced Benchmarks
Avg.
GSM8K
SVAMP
ASDiv
GSM-Hard
MATH
AIME24
AIME25
Minerva
Vanilla
92.42
84.00
95.98
51.78
81.20
73.33
66.67
33.46
72.36
EnvScaler
93.10
82.67
96.47
51.18
82.40
73.33
60.00
35.66
71.85
AgenticQwen
95.00
89.33
97.29
55.04
88.00
60.00
60.00
30.88
71.94
AgentSkiller
83.17
71.00
92.45
39.88
73.60
66.67
56.67
31.62
64.38
Klear-AgentForge
91.21
89.67
97.70
55.42
75.20
26.67
13.33
27.21
59.55
Table 1: Reasoning accuracy (%) across eight mathematical reasoning benchmarks. Vanilla denotes the initial Qwen3-8B model. Frozen Synthesizer retains online task generation but disables synthesis-policy updates. Avg. is the unweighted average across all benchmarks. Bold and underlined values indicate the best and second-best results, respectively.
Figure 2: Evolution of the synthesized curriculum during co-training. Tasks are ordered by generation time and divided into ten equally sized buckets. (a) Realized Reasoner outcomes, where gradient-bearing tasks have non-identical rollout rewards (equivalently, 0<SR<1 under our reward setting). (b) Evolution of synthesis signals, showing that task quality and novelty remain high while realized teachability declines. Generation order approximates training progression because exact iteration indices were not recorded.
Figure 3: Counterfactual enrichment analysis of synthesis signals. Tasks are reranked offline using different signals, and the proportion of gradient-bearing tasks in the top-ranked quartile is reported. Outcome-grounded teachability strongly enriches for gradient-bearing tasks, whereas static task-level signals provide substantially weaker separation.
Figure 4: Realized Reasoner outcomes under the Synthesizer’s ex-ante difficulty estimates. Harder labels shift tasks from uniformly correct toward uniformly incorrect outcomes, indicating coarse difficulty awareness. However, the gradient-bearing proportion remains nearly unchanged across difficulty categories, showing that nominal difficulty does not predict realized optimization utility.
Variant
Update
Wall
Grad. Rate
Tasks/Batch
Frozen
45.8 s
341.6 s
13.7%
4.38
SynCo
45.1 s
343.5 s
14.5%
4.64
Δ
−1.5%
+0.6%
+0.8 pp
+0.26
Table 3: Training cost and gradient utilization for Frozen Synthesizer and SynCo . Times are measured per training step, and productive rate denotes the fraction of gradient-bearing tasks.
Figure 5: Learner-relative comparison of synthetic data from SynCo , MetaMathQA, ScaleQuest-Math, and OpenMathInstruct-2 under the same Vanilla Qwen3-8B Reasoner and rollout protocol. (a) Distribution of realized rollout outcomes, highlighting the proportion of gradient-bearing tasks. (b) Full Reasoner success-rate distribution, showing how each data source is positioned relative to the learner’s capability boundary.
Figure 6: Reasoner reward trajectories under a frozen Synthesizer and joint Synthesizer–Reasoner co-training. Curves report 10-step moving averages. After comparable early progress, SynCo maintains a consistently higher reward trajectory.
Appendix figures & tables6 assets
Supplementary material from the paper’s appendix.
Appendix
Method
Primary Benchmarks
Advanced Benchmarks
Avg.
GSM8K
SVAMP
ASDiv
GSM-Hard
MATH-500
AIME24
AIME25
Minerva
Vanilla
92.34
86.25
96.63
52.01
82.60
63.33
56.67
31.25
70.14
EnvScaler
89.84
82.00
95.40
42.00
82.00
73.33
70.00
31.62
70.77
AgentSkiller
79.45
72.67
92.94
37.38
71.80
63.33
60.00
29.41
63.37
SynCo (Frozen Synth.)
93.10
92.00
97.78
61.33
86.20
66.67
60.67
31.99
73.72
SynCo
93.71
93.00
98.28
60.65
86.60
70.00
66.67
32.72
75.20
Appendix
Table 4: Reasoning accuracy (%) across eight mathematical reasoning benchmarks using Qwen3-4B. Vanilla denotes the initial Qwen3-4B model, while SynCo (Frozen Synth.) retains learner-conditioned online synthesis but disables Synthesizer updates. Avg. is the unweighted average across all benchmarks. Bold and underlined values indicate the best and second-best results, respectively.
(a) Realized Rollout Outcomes
Outcome Regime
Criterion
Number of Tasks
Percentage
All incorrect
SR=0
994
33.8%
Gradient-bearing
0<SR<1
475
16.2%
All correct
SR=1
1,471
50.0%
(b) Synthesis Signals vs. Realized Utility
Synthesis Signal
Correlation Target
Correlation
Observed Alignment
Appendix
Table 5: Mechanistic analysis of 2,940 verified tasks collected during co-training. Empirically, gradient-bearing groups coincide exactly with 0<SR<1 across the collected training data. Panel (a) summarizes realized rollout outcomes, Panel (b) relates synthesis signals to optimization utility, and Panel (c) compares the three major intended-difficulty groups with realized learnability; these groups cover 2,919 tasks.
Figure 7: Attribution of SynCo ’s terminal advantage over SynCo (Frozen Synth.). Retain measures additional solved examples among those initially solved by Qwen3-8B, whereas Learn measures gains on initially incorrect examples. Across all four benchmarks, most of the net improvement arises from the Learn subset, which contributes 45 of the 56 additional solved examples (80.4%).
Synthesis Signal
Mean
Std.
Corr. with Gradient-Bearing
Role
Interpretation within Verified Pool
Quality
0.962
0.030
+0.060
Reward gate
Nearly saturated after verification, with limited discrimination among valid tasks.
Reliability
0.965
0.096
+0.105
Reward gate
Ensures answer consistency and reliability, but weakly predicts realized gradient utility.
Report alignment
0.999
0.011
−0.021
Diagnostic
Nearly constant in the collected pool and weakly associated with realized utility.
Novelty
0.703
0.175
−0.041
Diagnostic
Varies substantially across tasks but does not identify gradient-bearing interactions.
Teachability
0.102
0.249
+0.934
Outcome utility
Strongly aligned with whether the task produces non-degenerate group-relative feedback.
Appendix
Table 6: Diagnostic statistics of synthesis signals over 2,940 verified tasks. Quality, reliability, and teachability constitute the gated reward used by SynCo , whereas novelty and report alignment are recorded as diagnostic signals. Correlation is measured against the binary gradient-bearing indicator defined by the Reasoner’s group-relative learning signal.
(a) Ex-Ante Difficulty vs. Realized Rollout Outcomes
Difficulty Label
Tasks
Mean SR
Gradient- Bearing
All Correct
All Incorrect
Learnable
630
0.625
16.8%
53.0%
30.2%
Mid-band
1,926
0.607
16.0%
51.1%
32.8%
Hard / Too hard
363
0.484
16.0%
39.9%
44.1%
Pool overall
2,940
0.594
16.2%
50.0%
33.8%
Appendix
Table 7: Calibration of the Synthesizer’s ex-ante difficulty estimates against realized Reasoner outcomes over 2,940 verified tasks. Panel (a) compares self-rated difficulty with the outcomes of K=4 Reasoner rollouts, while Panel (b) contrasts the ex-ante learnable label with outcome-grounded teachability as signals of realized gradient utility.
(a) Marginal Training Cost
Training Variant
Steps
Mean Update Time
Median Update Time
Median Wall Time
SynCo (Frozen Synth.)
100
45.8 s
45.0 s
341.6 s
SynCo
100
45.1 s
44.6 s
343.5 s
Relative difference
–
−1.5%
−0.9%
+0.6%
Appendix
Table 8: Detailed efficiency analysis of online Synthesizer–Reasoner co-training. Panel (a) compares per-step training cost, Panel (b) reports productive-task yield over the matched logging window, and Panel (c) summarizes task generation and verification during the 100-step SynCo run.
Large Language Models (LLMs) exhibit strong mathematical reasoning when trained on high-quality Chain-of-Thought (CoT) that articulates intermediate steps, yet costly CoT curation hinders further progress. While existing remedies such as distillation from stronger LLMs and self-synthesis based on test-time search alleviate this issue, they often suffer from diminishing returns or high computing overhead.In this work, we propose CoTEvol, a genetic evolutionary framework that casts CoT generation as a population-based search over reasoning trajectories.Candidate trajectories are iteratively evolved through reflective global crossover at the trajectory level and local mutation guided by uncertainty at the step level, enabling holistic recombination and fine-grained refinement. Lightweight, task-aware fitness functions are designed to guide the evolutionary process toward accurate and diverse reasoning. Empirically, CoTEvol improves correct-CoT synthesis success by over 30% and enhances structural diversity, with markedly improved efficiency. LLMs trained on these evolutionary CoT data achieve an average gain of 6.6% across eight math benchmarks, outperforming previous distillation and self-synthesis approaches. These results underscore the promise of evolutionary CoT synthesis as a scalable and effective method for mathematical reasoning tasks.
Zhuo Wang, Zhuo Zhang, Yafu Li +3
1Fudan University · 5Independent Researcher · 3The Chinese University of Hong Kong +3
Reinforcement learning for LLM agents is typically conducted on a static data distribution, which fails to adapt to the agent's evolving behavior and leads to poor coverage of complex environment interactions. To address these challenges, we propose CoEvolve, an agent-data mutual evolution framework that enables LLM agents to improve through closed-loop, interaction-driven training. Specifically, CoEvolve extracts feedback signals such as forgetting and uncertainty from rollout trajectories to identify failure-prone interaction patterns, and utilizes them to guide LLM-based task synthesis. The synthesized tasks are validated through environment interaction and utilized to update the data distribution, enabling joint adaptation of the agent and its data. Extensive experiments on AppWorld and BFCL across Qwen2.5-7B, Qwen3-4B, and Qwen3-30B-A3B demonstrate consistent and significant improvements over strong base models, yielding absolute gains of 19.43%, 15.58%, and 18.14%, respectively.
Large Language Model (LLM) Agents, often trained with Reinforcement Learning (RL), are constrained by a dependency on human-curated data, limiting scalability and tethering AI to human knowledge. Existing self-evolution frameworks offer an alternative but are typically restricted by the model's inherent capabilities and single-round interactions, hindering the development of complex curricula involving tool use or dynamic reasoning. We introduce Agent0, a fully autonomous framework that evolves high-performing agents without external data through multi-step co-evolution and seamless tool integration. Agent0 establishes a symbiotic competition between two agents initialized from the same base LLM: a curriculum agent that proposes increasingly challenging frontier tasks, and an executor agent that learns to solve them. We integrate external tools to enhance the executor's problem-solving capacity; this improvement, in turn, pressures the curriculum agent to construct more complex, tool-aware tasks. Through this iterative process, Agent0 establishes a self-reinforcing cycle that continuously produces high-quality curricula. Empirically, Agent0 substantially boosts reasoning capabilities, improving the Qwen3-8B-Base model by 18% on mathematical reasoning and 24% on general reasoning benchmarks. Code is available at https://github.com/aiming-lab/Agent0.
Peng Xia, Kaide Zeng, Jiaqi Liu +5
UNC-Chapel Hill · Salesforce Research · Stanford University