Self-evolving tool-integrated agents learn from tasks and feedback generated within their own training loop. A Curriculum Agent generates tasks, while an Executor Agent learns from self-consistency signals through reinforcement learning. However, relying solely on the current Executor for feedback has two limitations: group-relative advantages vanish under full consensus, while uncertainty-based curriculum rewards favor disagreement without showing whether the generated tasks support further learning. These limitations motivate an additional reference beyond the current Executor. We propose \textit{AnchorLoop}, which introduces a frozen copy of the previous iteration's Executor as a historical reference and reuses it on both sides of the training loop. For the Executor, the anchor provides a cross-reference advantage that evaluates current outputs against both current and historical majority answers. For the Curriculum, it provides an agreement-based reference based on differences in sampled majority agreement. Since the Executor and anchor have identical parameters during Curriculum training, this comparison serves as a proxy for task selection rather than evidence of inter-version improvement or correctness. Across 13 reasoning benchmarks, AnchorLoop improves over Agent0 by 2.5% on mathematical reasoning and 2.8% on general reasoning tasks. It also maintains higher effective-advantage variance and continues improving in later iterations as the unanchored baseline shows diminishing gains. These results demonstrate the benefit of introducing a lightweight historical reference into self-evolving tool-integrated agents without external task or answer supervision.
Figures & tables
Figure 1: Overview of AnchorLoop. The left panel illustrates an unanchored two-agent self-evolution loop that can plateau. At outer iteration t , AnchorLoop copies the previous Executor into a frozen anchor and reuses it on both sides of the loop. During Curriculum training, the current Executor and anchor have identical parameters but produce independent rollout groups for the agreement-based reference reward. All format-valid tasks form the same-iteration task pool Dt . During Executor training, the anchor remains fixed while the current Executor is updated. Current groups are resampled at every inner step, whereas one anchor group per task and Curriculum batch is cached and reused to construct the cross-reference advantage.
Mathematical reasoning
General reasoning
Model Name
AMC23
MATH
GSM8K
Minerva
Olympiad
AIME24
AIME25
AVG
SuperGPQA
MMLU-Pro
BBEH
GPQA-D
HumanEval
AGIEval
AVG
Qwen3-4B-Base
Base Model
✗
45.4
68.6
88.5
37.3
41.2
11.0
6.33
42.6
21.1
37.6
7.6
36.5
65.7
54.8
37.2
Base Model w/ tool
✓
45.9
72.7
88.8
38.1
42.4
12.5
7.86
44.0
25.8
43.6
8.5
37.2
67.1
55.4
39.6
+ Absolute Zero
✓
50.3
76.8
88.8
40.6
42.9
12.7
13.0
46.4
26.4
52.9
8.7
38.4
68.4
56.7
41.9
+ SPIRAL
✗
56.7
77.8
90.1
42.8
39.6
13.5
10.2
47.2
26.5
53.0
9.4
38.9
67.9
57.3
42.1
Table 1: Results on mathematical and general reasoning benchmarks. The two AVG columns summarize their respective benchmark groups. Best per backbone is in bold .
Figure 2: Training dynamics on Qwen3-4B-Base. (a) Math AVG across three outer iterations. (b) Within-group variance of the effective Executor advantage, Vadv , averaged over sampled groups. By iteration 3, it decreases to 0.07 for Agent0 but remains at 0.28 for AnchorLoop. (c) Fraction of Curriculum tasks with a positive sampled reference gap. An unresolved current–reference pair contributes zero. The fraction decreases from 54.7% initially to 16.1% for Agent0 and 32.9% for AnchorLoop by iteration 3.
Method
Math AVG
General AVG
Without tool access
Qwen3-4B (Base)
42.6
37.2
+ SPIRAL
47.2
42.1
+ R-Zero
48.7
42.8
With tool access
+ TIR
44.0
39.6
Table 2: Tool comparisons and anchor ablations on Qwen3-4B-Base. Agent0 and AnchorLoop are shared endpoints. Only the anchor variants use matched training and evaluation settings.
Figure 3: Anchor and Curriculum analyses on Qwen3-4B-Base. (a) Math AVG across outer iterations with an iterative anchor versus an anchor fixed to the base model. (b) Sensitivity over the historical-advantage weight α and Curriculum reference weight λ on a 4×4 grid; the default setting is (α,λ)=(0.6,0.4) . (c) Pass rate and average tool calls on 200 tasks sampled from each Curriculum stage. The iteration-1 Executor generates G=8 trajectories per task, and both statistics are computed from the same 1,600 trajectories per stage.
Appendix figures & tables10 assets
Supplementary material from the paper’s appendix.
Appendix
Group
Hyperparameter
Value
Outer loop
Outer iterations T
3
Group size G
8
Curriculum inner steps KC
32
Executor inner steps KE
128
Anchor side
Historical advantage weight α
0.6
Frontier reward weight λ
0.4
Appendix
Table 3: Training settings for both backbones. The base Curriculum reward weights are shared by baselines and AnchorLoop.
Backbone
Domain
Agent0
AnchorLoop
Qwen3-4B
Math
52.7±0.68
55.2±0.51
Qwen3-4B
General
45.1±0.57
47.9±0.44
MiMo-7B
Math
61.0±0.63
63.5±0.49
MiMo-7B
General
40.1±0.55
42.8±0.46
Appendix
Table 4: Run-level variability for the most recent baseline (Agent0) and AnchorLoop. Entries report mean ± sample standard deviation over five independent training runs.
Benchmark
Domain
Metric
Mathematical reasoning
AMC23
High school competition math
mean@32
MATH-500
Mixed math topics
pass@1
GSM8K
Grade school word problems
pass@1
Minerva
STEM quantitative reasoning
pass@1
OlympiadBench
Olympiad math
pass@1
Appendix
Table 5: Reasoning benchmarks used for evaluation.
Executor stage
Valid / candidate
Same majority
Zero historical term
Nonzero and collinear
Nonzero and non-collinear
Qwen3-4B-Base
Early
581/600
422(72.6%)
438(75.4%)
91(15.7%)
52(9.0%)
Middle
574/600
355(61.8%)
379(66.0%)
111(19.3%)
84(14.6%)
Late
586/600
347(59.2%)
369(63.0%)
103(17.6%)
114(19.5%)
MiMo-7B-Base
Early
578/600
414(71.6%)
430(74.4%)
96(16.6%)
52(9.0%)
Appendix
Table 6: Historical-advantage structure in matched records from five runs per backbone. Percentages use valid groups as the denominator. The final three columns partition the valid groups. The same-majority column may overlap with them.
Metric
Agent0
AnchorLoop (ours)
Time per outer iteration
20.5 h
24.1 h
GPU hours per outer iteration
164
193
Frozen checkpoint memory
0
8.1 GB
Anchor rollouts per task
0
G=8 (cached)
Overhead vs Agent0
–
+17.6%
Appendix
Table 7: Compute and memory comparison between Agent0 and AnchorLoop on Qwen3-4B-Base.
Refresh schedule
Best Math AVG
Every iteration (default)
56.3
Every two iterations
54.4
No refresh (fixed at base model)
53.3
Appendix
Table 8: Highest Math AVG among checkpoints evaluated within the first five outer iterations on Qwen3-4B-Base for each anchor refresh schedule.
Configuration
Wall time per outer iteration (h)
AnchorLoop without anchor caching
32.7
AnchorLoop with anchor caching
24.1
Agent0 (reference)
20.5
Appendix
Table 9: Wall time with and without anchor rollout caching on Qwen3-4B-Base, averaged over five runs.
Method (Iter)
AMC23
MATH
GSM8K
Minerva
Olympiad
AIME24
AIME25
AVG
Agent0
t=0
63.6
77.9
91.1
51.4
50.2
34.4
26.9
56.5
t=1
66.2
80.3
92.9
55.4
52.6
38.2
30.2
59.4
t=2
67.0
81.0
93.8
56.9
53.8
38.9
32.0
60.5
t=3
67.3
81.4
94.2
57.9
54.1
39.5
32.4
61.0
AnchorLoop (ours)
Appendix
Table 10: Per-benchmark math accuracy across outer iterations on MiMo-7B-Base. Iteration t=0 is the Base Model with tool. Iteration t=3 matches the endpoints in Table 1 . Best endpoint per benchmark is in bold .
Figure 4: Iteration dynamics on MiMo-7B-Base. (a) Math AVG across the three outer iterations. (b) The within-group variance of effective advantages, Vadv , averaged over sampled groups in the Executor stage. (c) The reported Curriculum frontier fraction, defined as the fraction of generated tasks with a positive sampled agreement gap under Appendix B.3 . An unresolved pair contributes zero.
Figure 5: Hyperparameter sensitivity on MiMo-7B-Base, the appendix counterpart of Figure 3 (b). Math AVG on a 4×4 grid over the historical-advantage weight α and the frontier reward weight λ . The peak at our default (α,λ)=(0.6,0.4) reaches 63.50 , matching Table 1 .