cs.CLSep 28, 2026

The Right Lesson at the Right Step: Deriving Control Updates for Self-Evolving Agents

Authors: Yunhe Su, ZiYi Dong, Tong Yu, Weijian Deng, Hao Li, Bowen Jiang, Pengxu Wei

Organizations: Sun Yat-sen University · eunomia-bpf community · Tsinghua Shenzhen International Graduate School, Tsinghua University · Nankai University · Department of Computer and Information Science, University of Pennsylvania, Philadelphia, PA, United States · Sun Yat-sen University, Peng Cheng Laboratory

Abstract

Self-evolving agents improve future behavior by reusing past experience, typically as global prompts, memories, or reflections. Yet these mechanisms rarely control where experience takes effect. In long tool-use workflows, the same lesson may correct one decision but distract another, making experience reuse a problem of localized control rather than memory alone. We introduce EvoCUE (Evolution through Control Updates from Evidence), a framework for learning reusable control-program updates from completed agent executions. EvoCUE represents the agent as an explicit state-machine controller, whose nodes perform model or tool calls and whose edges define where control passes next. This makes the workflow editable at precise locations, so each learned update can specify what to add, where it acts, and when it applies. From completed trajectories, EvoCUE uses residual goals and observed execution traces to propose localized instruction or skill edits. Each candidate is evaluated at the point where it would act by resuming the parent and edited controllers from the same checkpoint and comparing their final outcomes. Accepted edits are compiled with applicability rules, confirmed on held-out tasks, and inherited by later executions. We evaluate EvoCUE on long tool-use environments where learned conventions must reach the right execution step. From a minimal AppWorld controller without benchmark-specific onboarding instructions, EvoCUE learns the missing task-completion convention and substantially improves success on Test-Normal and Test-Challenge. On PAST-Bench office workflows, EvoCUE transfers organizational requirements from prior episodes to later tasks, improving task-execution quality. These results show that self-evolving agents should place experience inside the control flow, rather than only store it as text.

Figures & tables

Appendix figures & tables3 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. HarnessEvolve: Learning from Reference Trajectories for Reliable Agent Self-Evolution

    Sep 1, 2026Wen Jiang, Mingmin Chu, Yimeng Tian +6Self-Evolving AgentsSelf-Evolution

  2. EXG: Self-Evolving Agents with Experience Graphs

    May 18, 2026Yuxin Jin, Siyuan Zhang, Hanchen Wang +3Self-Evolving AgentsReasoning Benchmark

  3. Learning What to Remember and What to Internalize in LLM Self-Evolution via Adaptive Memory-Parameter Coordination

    Aug 2, 2026Tianyun Ji, Zhenya Huang, Jiayu Liu +3Self-Evolving AgentsSelf-Evolution