cs.AIOct 7, 2026

Self-Evolve With a Reference:Anchored Training of Tool-Integrated Agents

Authors: Wenjie Liao, Liangjie Zhao, Zehong Cao

Organizations: Waseda University · Adelaide University

Abstract

Self-evolving tool-integrated agents learn from tasks and feedback generated within their own training loop. A Curriculum Agent generates tasks, while an Executor Agent learns from self-consistency signals through reinforcement learning. However, relying solely on the current Executor for feedback has two limitations: group-relative advantages vanish under full consensus, while uncertainty-based curriculum rewards favor disagreement without showing whether the generated tasks support further learning. These limitations motivate an additional reference beyond the current Executor. We propose \textit{AnchorLoop}, which introduces a frozen copy of the previous iteration's Executor as a historical reference and reuses it on both sides of the training loop. For the Executor, the anchor provides a cross-reference advantage that evaluates current outputs against both current and historical majority answers. For the Curriculum, it provides an agreement-based reference based on differences in sampled majority agreement. Since the Executor and anchor have identical parameters during Curriculum training, this comparison serves as a proxy for task selection rather than evidence of inter-version improvement or correctness. Across 13 reasoning benchmarks, AnchorLoop improves over Agent0 by 2.5% on mathematical reasoning and 2.8% on general reasoning tasks. It also maintains higher effective-advantage variance and continues improving in later iterations as the unanchored baseline shows diminishing gains. These results demonstrate the benefit of introducing a lightweight historical reference into self-evolving tool-integrated agents without external task or answer supervision.

Figures & tables

Appendix figures & tables10 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Agent0: Unleashing Self-Evolving Agents from Zero Data via Tool-Integrated Reasoning

    Nov 20, 2025Peng Xia, Kaide Zeng, Jiaqi Liu +5Self-Evolving AgentsSelf-Evolution

  2. HarnessEvolve: Learning from Reference Trajectories for Reliable Agent Self-Evolution

    Sep 1, 2026Wen Jiang, Mingmin Chu, Yimeng Tian +6Self-Evolving AgentsSelf-Evolution

  3. UnifiedPlayers: Enhance Tool-Integrated Reasoning in Agentic Reinforcement Learning

    Sep 17, 2026Wenjie Liao, Liangjie Zhao, Zehong CaoAgentic Reinforcement LearningReasoning Benchmark