cs.AIOct 8, 2026

AgentEvolver: System-Wide Self-Evolution Through Task Execution

Authors: Wentao Zhang, Fuchao Yang, Yilei Zhao, Xinrun Wang, Bo An

Abstract

An agent can complete a task without improving how it works. Turning task experience into reusable capability requires connecting the changed component to its evaluation and subsequent use. We present AgentEvolver, a system for developing capabilities during task execution while keeping the foundation model fixed. Eight entity families expose reusable operations, methods, agents, control flow, interfaces, and supporting state to revision through a common versioned lifecycle. A shared Runtime coordinates ongoing work, while persistent planning and recoverable context preserve task direction and supporting evidence. We evaluate task outcomes on SWE-bench Pro Public and examine capability changes in six application cases. The team reports an 82.08% resolution rate with evolution, exceeding its reported baseline without evolution. The cases show retained capabilities entering later website, game, and research work, while also documenting incomplete objectives and an unsuccessful strategy. These findings distinguish improvement in a reusable component from success on the final task. AgentEvolver provides a concrete basis for studying capability accumulation through execution; independent-task transfer and total development cost remain open questions.

Explore similar work

CardsList
  1. EvoAgentBench: Benchmarking Agent Self-Evolution via Ability Transfer

    Jul 6, 2026Xingze Gao, Chuanrui Hu, Hongda Chen +9LLM Agent Self-ImprovementAI Agent Benchmarks

  2. Harness Updating Is Not Harness Benefit: Disentangling Evolution Capabilities in Self-Evolving LLM Agents

    May 28, 2026Minhua Lin, Juncheng Wu, Zijun Wang +14LLM Agent EvaluationAgent Harness Optimization

  3. EvolveNet: Collaborative Harness Evolution for Agent Self-Improvement

    Aug 5, 2026Jun Nie, Yonggang Zhang, Qianshu Cai +3Agent Harness OptimizationMulti-Agent Collaboration