Self-evolving LLM agents promise to improve autonomously through continual interaction and learning, reducing their dependence on manually curated supervision. Realizing this promise requires not only updating the agent, but also evolving its training experience as its capabilities change. However, most existing pipelines rely on static datasets or separately updated synthesis models, causing previously useful tasks to become trivial while overly difficult tasks remain uninformative. This growing mismatch between agent capability and training experience limits sustained self-improvement. To address this problem, we propose SynCo, an agentic data synthesis co-training framework for self-evolving LLMs based on multi-agent reinforcement learning. SynCo jointly optimizes two independently parameterized agents: a Synthesizer that constructs training tasks from the Reasoner's evolving capability state, and a Reasoner that learns from the resulting experience. Each synthesized task induces multiple Reasoner rollouts whose outcomes provide complementary rewards to both agents. Correctness feedback improves the Reasoner, while task quality, answer reliability, and outcome-grounded teachability guide the Synthesizer. Their updates are fed back into subsequent synthesis rounds, allowing the task-solving policy and its training distribution to evolve together. Extensive experiments across eight mathematical reasoning benchmarks demonstrate that SynCo substantially outperforms a broad range of existing synthetic-data methods and controlled baselines, achieving the strongest overall performance while deriving most of its gains from previously unsolved problems.
Figures & tables
Figure 1: Overview of SynCo , a closed-loop framework for self-evolving LLMs. (1) A Synthesizer first generates capability-guided and verifiable tasks from the current learning context. (2) The resulting interaction batch is then used to jointly update the Synthesizer and the Reasoner through online multi-agent co-training. (3) The updated agents in turn reshape the future curriculum and task distribution, forming an agent–data self-evolution loop.
Method
Primary Benchmarks
Advanced Benchmarks
Avg.
GSM8K
SVAMP
ASDiv
GSM-Hard
MATH
AIME24
AIME25
Minerva
Vanilla
92.42
84.00
95.98
51.78
81.20
73.33
66.67
33.46
72.36
EnvScaler
93.10
82.67
96.47
51.18
82.40
73.33
60.00
35.66
71.85
AgenticQwen
95.00
89.33
97.29
55.04
88.00
60.00
60.00
30.88
71.94
AgentSkiller
83.17
71.00
92.45
39.88
73.60
66.67
56.67
31.62
64.38
Klear-AgentForge
91.21
89.67
97.70
55.42
75.20
26.67
13.33
27.21
59.55
Table 1: Reasoning accuracy (%) across eight mathematical reasoning benchmarks. Vanilla denotes the initial Qwen3-8B model. Frozen Synthesizer retains online task generation but disables synthesis-policy updates. Avg. is the unweighted average across all benchmarks. Bold and underlined values indicate the best and second-best results, respectively.
Figure 2: Evolution of the synthesized curriculum during co-training. Tasks are ordered by generation time and divided into ten equally sized buckets. (a) Realized Reasoner outcomes, where gradient-bearing tasks have non-identical rollout rewards (equivalently, 0<SR<1 under our reward setting). (b) Evolution of synthesis signals, showing that task quality and novelty remain high while realized teachability declines. Generation order approximates training progression because exact iteration indices were not recorded.
Figure 3: Counterfactual enrichment analysis of synthesis signals. Tasks are reranked offline using different signals, and the proportion of gradient-bearing tasks in the top-ranked quartile is reported. Outcome-grounded teachability strongly enriches for gradient-bearing tasks, whereas static task-level signals provide substantially weaker separation.
Figure 4: Realized Reasoner outcomes under the Synthesizer’s ex-ante difficulty estimates. Harder labels shift tasks from uniformly correct toward uniformly incorrect outcomes, indicating coarse difficulty awareness. However, the gradient-bearing proportion remains nearly unchanged across difficulty categories, showing that nominal difficulty does not predict realized optimization utility.
Variant
Update
Wall
Grad. Rate
Tasks/Batch
Frozen
45.8 s
341.6 s
13.7%
4.38
SynCo
45.1 s
343.5 s
14.5%
4.64
Δ
−1.5%
+0.6%
+0.8 pp
+0.26
Table 3: Training cost and gradient utilization for Frozen Synthesizer and SynCo . Times are measured per training step, and productive rate denotes the fraction of gradient-bearing tasks.
Figure 5: Learner-relative comparison of synthetic data from SynCo , MetaMathQA, ScaleQuest-Math, and OpenMathInstruct-2 under the same Vanilla Qwen3-8B Reasoner and rollout protocol. (a) Distribution of realized rollout outcomes, highlighting the proportion of gradient-bearing tasks. (b) Full Reasoner success-rate distribution, showing how each data source is positioned relative to the learner’s capability boundary.
Figure 6: Reasoner reward trajectories under a frozen Synthesizer and joint Synthesizer–Reasoner co-training. Curves report 10-step moving averages. After comparable early progress, SynCo maintains a consistently higher reward trajectory.
Appendix figures & tables6 assets
Supplementary material from the paper’s appendix.
Appendix
Method
Primary Benchmarks
Advanced Benchmarks
Avg.
GSM8K
SVAMP
ASDiv
GSM-Hard
MATH-500
AIME24
AIME25
Minerva
Vanilla
92.34
86.25
96.63
52.01
82.60
63.33
56.67
31.25
70.14
EnvScaler
89.84
82.00
95.40
42.00
82.00
73.33
70.00
31.62
70.77
AgentSkiller
79.45
72.67
92.94
37.38
71.80
63.33
60.00
29.41
63.37
SynCo (Frozen Synth.)
93.10
92.00
97.78
61.33
86.20
66.67
60.67
31.99
73.72
SynCo
93.71
93.00
98.28
60.65
86.60
70.00
66.67
32.72
75.20
Appendix
Table 4: Reasoning accuracy (%) across eight mathematical reasoning benchmarks using Qwen3-4B. Vanilla denotes the initial Qwen3-4B model, while SynCo (Frozen Synth.) retains learner-conditioned online synthesis but disables Synthesizer updates. Avg. is the unweighted average across all benchmarks. Bold and underlined values indicate the best and second-best results, respectively.
(a) Realized Rollout Outcomes
Outcome Regime
Criterion
Number of Tasks
Percentage
All incorrect
SR=0
994
33.8%
Gradient-bearing
0<SR<1
475
16.2%
All correct
SR=1
1,471
50.0%
(b) Synthesis Signals vs. Realized Utility
Synthesis Signal
Correlation Target
Correlation
Observed Alignment
Appendix
Table 5: Mechanistic analysis of 2,940 verified tasks collected during co-training. Empirically, gradient-bearing groups coincide exactly with 0<SR<1 across the collected training data. Panel (a) summarizes realized rollout outcomes, Panel (b) relates synthesis signals to optimization utility, and Panel (c) compares the three major intended-difficulty groups with realized learnability; these groups cover 2,919 tasks.
Figure 7: Attribution of SynCo ’s terminal advantage over SynCo (Frozen Synth.). Retain measures additional solved examples among those initially solved by Qwen3-8B, whereas Learn measures gains on initially incorrect examples. Across all four benchmarks, most of the net improvement arises from the Learn subset, which contributes 45 of the 56 additional solved examples (80.4%).
Synthesis Signal
Mean
Std.
Corr. with Gradient-Bearing
Role
Interpretation within Verified Pool
Quality
0.962
0.030
+0.060
Reward gate
Nearly saturated after verification, with limited discrimination among valid tasks.
Reliability
0.965
0.096
+0.105
Reward gate
Ensures answer consistency and reliability, but weakly predicts realized gradient utility.
Report alignment
0.999
0.011
−0.021
Diagnostic
Nearly constant in the collected pool and weakly associated with realized utility.
Novelty
0.703
0.175
−0.041
Diagnostic
Varies substantially across tasks but does not identify gradient-bearing interactions.
Teachability
0.102
0.249
+0.934
Outcome utility
Strongly aligned with whether the task produces non-degenerate group-relative feedback.
Appendix
Table 6: Diagnostic statistics of synthesis signals over 2,940 verified tasks. Quality, reliability, and teachability constitute the gated reward used by SynCo , whereas novelty and report alignment are recorded as diagnostic signals. Correlation is measured against the binary gradient-bearing indicator defined by the Reasoner’s group-relative learning signal.
(a) Ex-Ante Difficulty vs. Realized Rollout Outcomes
Difficulty Label
Tasks
Mean SR
Gradient- Bearing
All Correct
All Incorrect
Learnable
630
0.625
16.8%
53.0%
30.2%
Mid-band
1,926
0.607
16.0%
51.1%
32.8%
Hard / Too hard
363
0.484
16.0%
39.9%
44.1%
Pool overall
2,940
0.594
16.2%
50.0%
33.8%
Appendix
Table 7: Calibration of the Synthesizer’s ex-ante difficulty estimates against realized Reasoner outcomes over 2,940 verified tasks. Panel (a) compares self-rated difficulty with the outcomes of K=4 Reasoner rollouts, while Panel (b) contrasts the ex-ante learnable label with outcome-grounded teachability as signals of realized gradient utility.
(a) Marginal Training Cost
Training Variant
Steps
Mean Update Time
Median Update Time
Median Wall Time
SynCo (Frozen Synth.)
100
45.8 s
45.0 s
341.6 s
SynCo
100
45.1 s
44.6 s
343.5 s
Relative difference
–
−1.5%
−0.9%
+0.6%
Appendix
Table 8: Detailed efficiency analysis of online Synthesizer–Reasoner co-training. Panel (a) compares per-step training cost, Panel (b) reports productive-task yield over the matched logging window, and Panel (c) summarizes task generation and verification during the 100-step SynCo run.