Agent evolvers automate the design of the prompts, skills and workflows around language model agents, yet the optimization process they follow is still designed by hand: a fixed search loop decides how candidates are evaluated, which are kept and when the search stops. We propose FREEEVOLVE, which automates this process as well. An environment specifies the goal, target agent, evaluator, data and resource limits; within these limits, the evolver itself decides what to test, how much evidence to collect, which candidates to pursue and when to stop. These decisions follow an editable evolution skill, which we improve through meta-evolution by scoring each candidate skill on the fresh target agent it produces. The optimization process thus becomes a capability learned from experience rather than a loop engineered in advance. On tau3-bench, ARC-AGI-2, ARC-AGI-3 and Terminal-Bench 2.1, FREEEVOLVE controls the evolution campaign by itself, yet improves the primary held-out metric by 13.6 points on average and matches or exceeds hand-designed evolvers. The learned process keeps improving with experience: meta-evolved skills add 6.9 points over the seed skill on fresh target agents, demonstrating transferability across environments.
Figures & tables
Figure 1: From prescribed loops to an automated, learned process. (a) Prescribed loop. A prescribed loop fixes all campaign decisions and their order before evolution begins. (b) Campaign autonomy. FreeEvolve transfers control of these decisions to the evolver, allowing it to adapt the campaign as new evidence becomes available. (c) Meta-evolution. Campaign autonomy alone provides limited improvements. With the hand-written seed skill, the evolver improves a fresh target agent by at most 1.1 points. Meta-evolution evaluates candidate skills through complete improvement campaigns and the target agents they produce. Once frozen and installed on fresh target agents, the learned skills improve held-out performance by 2.2 to 13.8 points (see Figure 3 ).
Figure 2: The interface at both levels. (a) The environment fixes the goal, data, editable target, execution function, and valid requests; the evolver is invoked once. Each request (x,u) chooses a candidate and how to run it (which tasks, how many repeats, what concurrency and timeout). The returned artifacts update the campaign state, which steers the next request, and the evolver returns x⋆ when it decides to stop. (b) Meta-evolution applies the same Evolve with the skill as the target: one execution of Rm runs complete worker campaigns, and a candidate skill is judged only by the workers they produce.
τ3 -bench
ARC-AGI-2
ARC-AGI-3
Terminal-Bench 2.1
Method
Mean
Pass@3
Mean
Pass@3
RHAE
Mean
Pass@3
Base agent
33.3 ±5.3
48.3
55.4 ±4.0
75.0
61.0 ±1.1
80.9 ±5.0
88.5
GEPA ( 2025 )
35.6 ±2.5
51.7
58.0 ±5.7
76.4
65.8 ±2.6
84.7 ±2.4
91.8
OpenEvolve ( 2025 )
35.6 ±2.4
55.2
58.6 ±5.0
79.1
63.4 ±1.5
82.5 ±1.8
90.2
EvoX ( 2026a )
36.8 ±3.5
58.6
64.0 ±6.4
80.0
65.1 ±3.4
82.0 ±1.8
90.2
HyperAgents ( 2026b )
36.8 ±5.1
58.6
59.6 ±4.2
77.8
72.1 ±2.8
82.5 ±2.1
90.2
Table 1: FreeEvolve matches or exceeds hand-designed evolvers. Held-out performance (%) of the worker each method returns; means ± s.d.; RHAE: Relative Human Action Efficiency. EvoX and HyperAgents also meta-evolve their search procedure. The last row is its relative gain over the base agent.
Figure 3: Meta-evolved skills make stronger evolvers on fresh workers. Each skill is frozen and installed on the same fresh base worker with the same budget. Δ : meta-evolved minus seed.
Target: τ3 -bench
Target: ARC-AGI-2
Target: ARC-AGI-3
Target: TB 2.1
Meta-evolution environment(s)
Mean
Pass@3
Mean
Pass@3
RHAE
Mean
Pass@3
Minimal seed
33.3 ±5.3
48.3
56.5 ±1.4
74.8
61.7 ±4.3
81.4 ±2.5
86.9
τ3 -bench
47.1 ±8.7
62.1
64.3 ±2.8
83.8
71.0 ±6.4
76.5 ±7.4
90.2
ARC-AGI-2
44.8 ±3.4
55.2
59.8 ±1.9
81.1
62.5 ±3.6
74.3 ±5.0
85.2
τ3 -bench + TB2.1
46.0 ±8.7
51.7
58.0 ±1.9
81.1
66.9 ±4.8
82.5 ±4.1
91.8
Table 2: Learned evolution skills transfer across environments. Each row is a skill meta-evolved on the listed environment(s), frozen, and installed on a fresh worker in each target; shading marks improvement over the seed. Single-source skills transfer unevenly and lower the Terminal-Bench mean; meta-evolving on τ3 -bench and Terminal-Bench 2.1 (TB2.1) jointly improves all metrics.
Figure 4: Evolved harnesses transfer to cheaper worker models, and FreeEvolve learns a strong and efficient skill. (a) Each method evolves the harness with the native model; the frozen harness then runs unchanged with Sonnet 4.6 and Haiku 4.5. FreeEvolve is the strongest harness in all six settings and keeps a 5.3 -point average gain over the base agent on the four non-native ones. (b) Best-so-far ARC-AGI-3 training RHAE of the worker each method’s skill produces, versus the total tokens used to produce that skill.
Evolver model
Mean (%)
Pass@3
No meta-evolution
33.3 ±5.3
48.3
Haiku 4.5
32.2 ±8.7
48.3
Sonnet 4.6
47.1 ±4.0
62.1
Opus 4.6
40.2 ±2.0
55.2
Opus 4.7
49.4 ±8.0
69.0
Opus 4.8
51.7 ±6.0
69.0
Table 3: Evolver-model ablation on τ3 -bench .
Campaign control
Mean (%)
Tokens (B)
Controlled u
47.1
1.19
Full set
39.1
0.89
Concurrency =24
43.7
1.78
No timeout
44.8
1.53
Table 4: Fixed-decision ablations on τ3 -bench .
Figure 5: The evolver reorganizes evaluation and search as evidence arrives. One ARC-AGI-2 worker-evolution campaign under a frozen meta-evolved skill: the evolution-set score of every evaluation against elapsed hours (markers as in the legend). Annotations mark five decisions the evolver made on its own: undoing a three-edit regression while keeping the useful helper library, probing a 2% subset to show that a truncation came from the agent’s own prompt rather than the platform, merging the best code with a fix from a weaker run, repeating a promising version to confirm its gain, and branching an ancestor into four variants. The star marks the harness it selects.
Appendix figures & tables10 assets
Supplementary material from the paper’s appendix.
Appendix
Benchmark
Dataset and primary metric
Evolution / held-out protocol
Base agent
τ3 -bench
97 interactive banking_knowledge conversations with tools and a domain policy; mean task pass rate and pass@3.
68 evolution tasks and 29 disjoint held-out tasks.
Vendored τ3 customer-service agent with Claude Opus 4.8.
ARC-AGI-2
Few-shot visual grid-transformation tasks that test abstract rule induction; mean task accuracy and pass@3.
56 evolution tasks and 111 disjoint validation tasks.
Base ARC solver backed by Claude 4.7 Opus.
ARC-AGI-3
Interactive abstract-reasoning games; Relative Human Action Efficiency (RHAE) score.
17 evolution games and 8 disjoint held-out games.
Base ARC-AGI-3 agent backed by Claude Opus 5.
Terminal-Bench 2.1
Containerized command-line tasks covering software, systems, security, and scientific workflows; mean task pass rate and pass@3.
Evolution tasks are hidden from final evaluation; 61 held-out tasks.
Unmodified deep agent with Claude Opus 4.8.
Appendix
Table 5: Benchmark datasets, metrics, evaluation protocols, and base agents.
Figure 6: Campaign rules added to the seed skill by meta-evolution.
Dataset
Metric
Base worker
Seed skill
No-data LLM
Meta-evolved
Δ
τ3 -bench
Mean
33.3 ±5.3
33.3 ±5.3
37.9 ±6.0
47.1 ±8.7
+13.8
Pass@3
48.3
48.3
58.6
62.1
+13.8
ARC-AGI-2
Mean
55.4 ±4.0
56.5 ±1.4
56.8 ±3.1
59.8 ±1.9
+3.3
Pass@3
75.0
74.8
74.8
81.1
+6.3
ARC-AGI-3
RHAE
61.0 ±1.1
61.7 ±4.3
64.5 ±5.1
70.3 ±3.2
+8.6
Terminal-Bench 2.1
Mean
80.9 ±5.0
81.4 ±2.5
78.7 ±3.3
83.1 ±4.1
+1.7
Appendix
Table 6: Fresh-worker evaluation of seed, no-data, and meta-evolved evolvers.
Meta-evolution model
Evolution
Held-out mean
Pass@3
No meta-evolution
35.0
33.3 ±5.3
48.3
Haiku 4.5
52.2
32.2 ±8.7
48.3
Sonnet 4.6
56.4
47.1 ±4.0
62.1
Opus 4.6
48.0
40.2 ±2.0
55.2
Opus 4.7
59.5
49.4 ±8.0
69.0
Opus 4.8
51.5
51.7 ±6.0
69.0
Appendix
Table 7: Effect of evolver-model capability on meta-evolution for τ3 -bench.
Evolver-controlled u
Always full set
Concurrency =24
No task timeout
Held-out mean
47.1 ±8.7
39.1 ±4.3
43.7 ±4.3
44.8 ±3.5
Pass@3
62.1
58.6
58.6
55.2
Total tokens
1.19B
0.89B
1.78B
1.53B
Appendix
Table 8: Performance and total token use when individual campaign decisions are fixed on τ3 -bench.
Method
τ3 -bench
ARC-AGI-2
ARC-AGI-3
TB 2.1
GEPA
11.0
11.9
30.1
32.1
OpenEvolve
9.3
16.5
36.0
25.7
EvoX
6.9
13.0
33.5
25.8
HyperAgents
11.9
13.4
23.8
29.7
FreeEvolve meta-evolution
24.0
24.0
24.0
24.0
Appendix
Table 9: Wall-clock hours by method and benchmark. The FreeEvolve row reports its 24-hour outer-campaign budget.
Native source model
Sonnet 4.6
Haiku 4.5
Dataset
Evolution method
Mean
Pass@3
Mean
Pass@3
Mean
Pass@3
τ3 -bench
Base agent
33.3 ±5.3
48.3
40.2 ±8.0
55.2
17.2 ±6.0
34.5
GEPA
35.6 ±4.0
51.7
36.8 ±8.0
51.7
9.2 ±4.0
20.7
OpenEvolve
35.6 ±4.0
55.2
35.6 ±5.3
62.1
14.9 ±5.3
31.0
EvoX
36.8 ±2.0
58.6
39.1 ±8.0
44.8
11.5 ±5.3
24.1
HyperAgents
36.8 ±5.3
58.6
35.6 ±4.0
48.3
14.9 ±3.0
31.0
Appendix
Table 10: Complete cross-model transfer results for frozen worker harnesses.
Question
Evidence
Q1: Can a learned autonomous campaign achieve strong performance without a prescribed loop?
FreeEvolve improves the primary held-out metric by 13.6 points on average over the base worker, outperforms the expert-designed evolvers on three of four benchmarks, and matches or outperforms the meta-evolution baselines EvoX and HyperAgents. On Terminal-Bench 2.1 it has the best pass@3 but trails GEPA in mean score (Table 1 ).
Q2: Does meta-evolution turn campaign autonomy into a stronger evolver?
On fresh workers, the seed skill gains only 0.6 points on average across the four mean metrics. The frozen meta-evolved skill adds a further 6.9 points, improves every available pass@3 metric, and outperforms the no-data LLM rewrite in every environment (Figure 3 ).
Q3: Does the resulting evolution skill generalize beyond its source environment and model?
Frozen skills transfer positively on most target metrics, but both single-source skills reduce the Terminal-Bench 2.1 mean; joint meta-evolution on τ3 -bench and Terminal-Bench 2.1 improves every target metric over the seed (Table 2 ). Evolved worker harnesses also transfer across models: run unchanged with Sonnet 4.6 and Haiku 4.5, FreeEvolve keeps a 5.3 -point average gain over the base harness (Figure 4 a).
Q4: How does the evolver use campaign autonomy as evolution unfolds?
Fixing any single evaluation decision (always the full set, concurrency =24 , or no timeout) yields a weaker worker, and the last two also use more tokens (Table 4 ). The ARC-AGI-2 trace shows the evolver changing evaluation resolution, opening branches, combining discoveries, and selecting the final worker as evidence accumulates (Figure 5 ); meta-evolved skills turn the seed’s dispositions into operational rules (Figure 6 ).
Q5: What model capability is required to exercise campaign autonomy effectively?
With Haiku 4.5 as the evolver, meta-evolution improves neither held-out metric. From Sonnet 4.6 onward it consistently improves held-out performance, with the strongest Opus models producing the best workers, although Sonnet 4.6 outperforms Opus 4.6 (Table 3 ).
Appendix
Table 11: Experimental questions Q1–Q5 and the main-text evidence that answers each.
Contract field
Worker environment: ARC-AGI-3
Meta-evolver environment: meta_chain
Goal
Improve the game-playing agent so it completes more ARC-AGI-3 levels with fewer actions.
Improve the evolution skill itself, and use it to produce strong agents in every installed inner environment.
Editable target
agent/arc3_agent/ : the worker’s harness and game-playing loop.
target/SKILL.md : the instructions that control a complete worker-evolution campaign.
Held fixed
Model, reasoning effort, token budget, game engine, games, scoring rule, 2,000-action limit, and 7,200-second attempt limit.
Inner environments and their evaluators. The seed’s dispositions are preserved while its operational method is evolved.
One evaluation
Run a worker version on selected games and attempts: ./run.sh --artifacts <dir> [--tasks ...] [--attempts ...] .
Install the candidate skill as an inner evolver and launch a complete worker-evolution run: ./run.sh --artifacts <dir> --inner-env arc_agi_3 .
Evidence
A scalar Relative Human Action Efficiency (RHAE) mean plus levels completed, actions, errors, per-game results, action traces, and agent transcripts.
No single scalar. The outer evolver reads the inner run’s evaluations, candidate workers, final absolute quality, cost, and campaign trace.
Evaluation controls
Choose games, attempts, concurrency, and whole-evaluation timeout.
Choose the inner environment and run candidate skills concurrently.
Appendix
Table 12: ARC-AGI-3 worker and meta-evolver environments. The shared contract fields hide different targets and units of evaluation.
Fixed by the canonical seed
Left open to the evolver
Environment description as source of truth; edit only the declared target
Which tasks or failure clusters to evaluate next
Preserve raw runs and artifacts; make claims proportional to evidence
Evaluation configuration, replicate allocation, and numerical promotion tests
Use the available time; continue from the best-supported target
Branching, parallelism, backtracking, composition, and search-layer changes
Return the strongest target under best/ with an evidence summary
Internal candidate organization and the path used to reach the nominee
Appendix
Table 13: What the canonical seed fixes and what remains available for meta-evolution to learn.
Language agents are expected to solve increasingly complex tasks, creating a growing need for continual improvement. One promising approach is to evolve the agent harness, the software that governs tool use, memory management, and task execution, while keeping the underlying language model fixed. Recent methods automate this process by using a meta coding agent to modify the harness based on execution feedback. However, relying on that agent's existing knowledge and observed failures can restrict exploration and make adaptation reactive. Inspired by how human experts learn from the research literature for new solutions, we introduce ScholarEvolve, a framework that automatically draws on state-of-the-art research to guide harness evolution. ScholarEvolve organizes the harness evolution directions into functional modules and uses topic modeling to identify distinct improvement strategies for each module. It implements these strategies and evaluates their combinations to improve task performance. Moreover, the framework is designed to incorporate new publications over time, allowing research advances to drive proactive lifelong evolution. Experiments demonstrate improvements on AppWorld and Tau2-Bench. ScholarEvolve raises Qwen3.5-27B task goal completion from 49.6% to 63.6% on AppWorld Challenge, and raises GPT-5.4-mini pass@1 from 72.7% to 81.9% on Tau2-Bench Telecom.
Jingbo Yang, Kwei-Herng Lai, Xiaowen Wang +3
University of California, Santa Barbara · Microsoft
Most agents today ``self-evolve'' by following rewards and rules defined by humans. However, this process remains fundamentally dependent on external supervision; without human guidance, the evolution stops. In this work, we train agents to possess an intrinsic meta-evolution capability to spontaneously learn about unseen environments prior to task execution. To instill this ability, we design an outcome-based reward mechanism that measures how much an agent's self-generated world knowledge improves its success rate on downstream tasks. This reward signal is used exclusively during the training phase to teach the model how to explore and summarize effectively. At inference time, the agent requires no external rewards or human instructions. It spontaneously performs native self-evolution to adapt to unknown environments using its internal parameters. When applied to Qwen3-30B and Seed-OSS-36B, this shift to native evolution yields a 20% performance increase on WebVoyager and WebWalker. Most strikingly, the generated world knowledge even enables a compact 14B Qwen3 model to outperform the unassisted Gemini-2.5-Flash, establishing a new paradigm for truly evolving agents.
Qifan Zhang, Dongyang Ma, Tianqing Fang +5
1Tencent · 2The Hong Kong University of Science and Technology (Guangzhou)
An LLM agent's capability depends not only on model weights but on its harness: prompts, tools, skills, and control flow. Self-improvement loops already edit harnesses, yet single-lineage search is path-dependent and local wins often regress other tasks. We introduce DarwinX, which treats self-evolution as selection over a population of harnesses with the model frozen: a preserve-and-extend contract admits only variants that extend coverage without regressing, an archive keeps alternative lineages for recombination, and failure-, teacher-, and self-derived evidence share one edit interface. Fitness comes from each benchmark's own verifier: no gold solutions, no hand-picked winners. Across four benchmarks that progressively separate the evolution signal from the test, one loop adds about 17 points on average: Terminal-Bench 2.1 rises +7.7 to 83.2% on a matched base and to the verified frontier at 84.7% on a stronger one; TerminalWorld's held-out split reaches 68.3%, ahead of every off-the-shelf agent; WebArena-Infinity real-task pass@1 rises from 43.5% to 93.0% audit-clean; and a Terminal-Bench 2.1 harness transfers unchanged to SWE-bench Verified. What evolves is general agent competence, not benchmark-specific patches, so it survives changes of task, verifier, and base model. A frozen model need not be a fixed agent: harness selection turns evaluation compute into durable capability.