FreeEvolve: Learning to Evolve Beyond Fixed Loops
Organizations: AWS AI Labs · Georgia Institute of Technology
Abstract
Agent evolvers automate the design of the prompts, skills and workflows around language model agents, yet the optimization process they follow is still designed by hand: a fixed search loop decides how candidates are evaluated, which are kept and when the search stops. We propose FREEEVOLVE, which automates this process as well. An environment specifies the goal, target agent, evaluator, data and resource limits; within these limits, the evolver itself decides what to test, how much evidence to collect, which candidates to pursue and when to stop. These decisions follow an editable evolution skill, which we improve through meta-evolution by scoring each candidate skill on the fresh target agent it produces. The optimization process thus becomes a capability learned from experience rather than a loop engineered in advance. On tau3-bench, ARC-AGI-2, ARC-AGI-3 and Terminal-Bench 2.1, FREEEVOLVE controls the evolution campaign by itself, yet improves the primary held-out metric by 13.6 points on average and matches or exceeds hand-designed evolvers. The learned process keeps improving with experience: meta-evolved skills add 6.9 points over the seed skill on fresh target agents, demonstrating transferability across environments.
Figures & tables
| -bench | ARC-AGI-2 | ARC-AGI-3 | Terminal-Bench 2.1 | ||||
| Method | Mean | Pass@3 | Mean | Pass@3 | RHAE | Mean | Pass@3 |
| Base agent | 33.3 | 48.3 | 55.4 | 75.0 | 61.0 | 80.9 | 88.5 |
| GEPA ( 2025 ) | 35.6 | 51.7 | 58.0 | 76.4 | 65.8 | 84.7 | 91.8 |
| OpenEvolve ( 2025 ) | 35.6 | 55.2 | 58.6 | 79.1 | 63.4 | 82.5 | 90.2 |
| EvoX ( 2026a ) | 36.8 | 58.6 | 64.0 | 80.0 | 65.1 | 82.0 | 90.2 |
| HyperAgents ( 2026b ) | 36.8 | 58.6 | 59.6 | 77.8 | 72.1 | 82.5 | 90.2 |
| Target: -bench | Target: ARC-AGI-2 | Target: ARC-AGI-3 | Target: TB 2.1 | ||||
|---|---|---|---|---|---|---|---|
| Meta-evolution environment(s) | Mean | Pass@3 | Mean | Pass@3 | RHAE | Mean | Pass@3 |
| Minimal seed | 33.3 | 48.3 | 56.5 | 74.8 | 61.7 | 81.4 | 86.9 |
| -bench | 47.1 | 62.1 | 64.3 | 83.8 | 71.0 | 76.5 | 90.2 |
| ARC-AGI-2 | 44.8 | 55.2 | 59.8 | 81.1 | 62.5 | 74.3 | 85.2 |
| -bench + TB2.1 | 46.0 | 51.7 | 58.0 | 81.1 | 66.9 | 82.5 | 91.8 |
| Evolver model | Mean (%) | Pass@3 |
|---|---|---|
| No meta-evolution | 33.3 | 48.3 |
| Haiku 4.5 | 32.2 | 48.3 |
| Sonnet 4.6 | 47.1 | 62.1 |
| Opus 4.6 | 40.2 | 55.2 |
| Opus 4.7 | 49.4 | 69.0 |
| Opus 4.8 | 51.7 | 69.0 |
| Campaign control | Mean (%) | Tokens (B) |
|---|---|---|
| Controlled | 47.1 | 1.19 |
| Full set | 39.1 | 0.89 |
| Concurrency | 43.7 | 1.78 |
| No timeout | 44.8 | 1.53 |
Appendix figures & tables10 assets
Supplementary material from the paper’s appendix.
Appendix
| Benchmark | Dataset and primary metric | Evolution / held-out protocol | Base agent |
|---|---|---|---|
| -bench | 97 interactive banking_knowledge conversations with tools and a domain policy; mean task pass rate and pass@3. | 68 evolution tasks and 29 disjoint held-out tasks. | Vendored customer-service agent with Claude Opus 4.8. |
| ARC-AGI-2 | Few-shot visual grid-transformation tasks that test abstract rule induction; mean task accuracy and pass@3. | 56 evolution tasks and 111 disjoint validation tasks. | Base ARC solver backed by Claude 4.7 Opus. |
| ARC-AGI-3 | Interactive abstract-reasoning games; Relative Human Action Efficiency (RHAE) score. | 17 evolution games and 8 disjoint held-out games. | Base ARC-AGI-3 agent backed by Claude Opus 5. |
| Terminal-Bench 2.1 | Containerized command-line tasks covering software, systems, security, and scientific workflows; mean task pass rate and pass@3. | Evolution tasks are hidden from final evaluation; 61 held-out tasks. | Unmodified deep agent with Claude Opus 4.8. |
| Dataset | Metric | Base worker | Seed skill | No-data LLM | Meta-evolved | |
|---|---|---|---|---|---|---|
| -bench | Mean | 33.3 | 33.3 | 37.9 | 47.1 | |
| Pass@3 | 48.3 | 48.3 | 58.6 | 62.1 | ||
| ARC-AGI-2 | Mean | 55.4 | 56.5 | 56.8 | 59.8 | |
| Pass@3 | 75.0 | 74.8 | 74.8 | 81.1 | ||
| ARC-AGI-3 | RHAE | 61.0 | 61.7 | 64.5 | 70.3 | |
| Terminal-Bench 2.1 | Mean | 80.9 | 81.4 | 78.7 | 83.1 |
| Meta-evolution model | Evolution | Held-out mean | Pass@3 |
|---|---|---|---|
| No meta-evolution | 35.0 | 33.3 | 48.3 |
| Haiku 4.5 | 52.2 | 32.2 | 48.3 |
| Sonnet 4.6 | 56.4 | 47.1 | 62.1 |
| Opus 4.6 | 48.0 | 40.2 | 55.2 |
| Opus 4.7 | 59.5 | 49.4 | 69.0 |
| Opus 4.8 | 51.5 | 51.7 | 69.0 |
| Evolver-controlled | Always full set | Concurrency | No task timeout | |
|---|---|---|---|---|
| Held-out mean | 47.1 | 39.1 | 43.7 | 44.8 |
| Pass@3 | 62.1 | 58.6 | 58.6 | 55.2 |
| Total tokens | 1.19B | 0.89B | 1.78B | 1.53B |
| Method | -bench | ARC-AGI-2 | ARC-AGI-3 | TB 2.1 |
|---|---|---|---|---|
| GEPA | 11.0 | 11.9 | 30.1 | 32.1 |
| OpenEvolve | 9.3 | 16.5 | 36.0 | 25.7 |
| EvoX | 6.9 | 13.0 | 33.5 | 25.8 |
| HyperAgents | 11.9 | 13.4 | 23.8 | 29.7 |
| FreeEvolve meta-evolution | 24.0 | 24.0 | 24.0 | 24.0 |
| Native source model | Sonnet 4.6 | Haiku 4.5 | |||||
|---|---|---|---|---|---|---|---|
| Dataset | Evolution method | Mean | Pass@3 | Mean | Pass@3 | Mean | Pass@3 |
| -bench | Base agent | 33.3 | 48.3 | 40.2 | 55.2 | 17.2 | 34.5 |
| GEPA | 35.6 | 51.7 | 36.8 | 51.7 | 9.2 | 20.7 | |
| OpenEvolve | 35.6 | 55.2 | 35.6 | 62.1 | 14.9 | 31.0 | |
| EvoX | 36.8 | 58.6 | 39.1 | 44.8 | 11.5 | 24.1 | |
| HyperAgents | 36.8 | 58.6 | 35.6 | 48.3 | 14.9 | 31.0 | |
| Question | Evidence |
|---|---|
| Q1: Can a learned autonomous campaign achieve strong performance without a prescribed loop? | FreeEvolve improves the primary held-out metric by points on average over the base worker, outperforms the expert-designed evolvers on three of four benchmarks, and matches or outperforms the meta-evolution baselines EvoX and HyperAgents. On Terminal-Bench 2.1 it has the best pass@3 but trails GEPA in mean score (Table 1 ). |
| Q2: Does meta-evolution turn campaign autonomy into a stronger evolver? | On fresh workers, the seed skill gains only points on average across the four mean metrics. The frozen meta-evolved skill adds a further points, improves every available pass@3 metric, and outperforms the no-data LLM rewrite in every environment (Figure 3 ). |
| Q3: Does the resulting evolution skill generalize beyond its source environment and model? | Frozen skills transfer positively on most target metrics, but both single-source skills reduce the Terminal-Bench 2.1 mean; joint meta-evolution on -bench and Terminal-Bench 2.1 improves every target metric over the seed (Table 2 ). Evolved worker harnesses also transfer across models: run unchanged with Sonnet 4.6 and Haiku 4.5, FreeEvolve keeps a -point average gain over the base harness (Figure 4 a). |
| Q4: How does the evolver use campaign autonomy as evolution unfolds? | Fixing any single evaluation decision (always the full set, concurrency , or no timeout) yields a weaker worker, and the last two also use more tokens (Table 4 ). The ARC-AGI-2 trace shows the evolver changing evaluation resolution, opening branches, combining discoveries, and selecting the final worker as evidence accumulates (Figure 5 ); meta-evolved skills turn the seed’s dispositions into operational rules (Figure 6 ). |
| Q5: What model capability is required to exercise campaign autonomy effectively? | With Haiku 4.5 as the evolver, meta-evolution improves neither held-out metric. From Sonnet 4.6 onward it consistently improves held-out performance, with the strongest Opus models producing the best workers, although Sonnet 4.6 outperforms Opus 4.6 (Table 3 ). |
| Contract field | Worker environment: ARC-AGI-3 | Meta-evolver environment: meta_chain |
|---|---|---|
| Goal | Improve the game-playing agent so it completes more ARC-AGI-3 levels with fewer actions. | Improve the evolution skill itself, and use it to produce strong agents in every installed inner environment. |
| Editable target | agent/arc3_agent/ : the worker’s harness and game-playing loop. | target/SKILL.md : the instructions that control a complete worker-evolution campaign. |
| Held fixed | Model, reasoning effort, token budget, game engine, games, scoring rule, 2,000-action limit, and 7,200-second attempt limit. | Inner environments and their evaluators. The seed’s dispositions are preserved while its operational method is evolved. |
| One evaluation | Run a worker version on selected games and attempts: ./run.sh --artifacts <dir> [--tasks ...] [--attempts ...] . | Install the candidate skill as an inner evolver and launch a complete worker-evolution run: ./run.sh --artifacts <dir> --inner-env arc_agi_3 . |
| Evidence | A scalar Relative Human Action Efficiency (RHAE) mean plus levels completed, actions, errors, per-game results, action traces, and agent transcripts. | No single scalar. The outer evolver reads the inner run’s evaluations, candidate workers, final absolute quality, cost, and campaign trace. |
| Evaluation controls | Choose games, attempts, concurrency, and whole-evaluation timeout. | Choose the inner environment and run candidate skills concurrently. |
| Fixed by the canonical seed | Left open to the evolver |
|---|---|
| Environment description as source of truth; edit only the declared target | Which tasks or failure clusters to evaluate next |
| Preserve raw runs and artifacts; make claims proportional to evidence | Evaluation configuration, replicate allocation, and numerical promotion tests |
| Use the available time; continue from the best-supported target | Branching, parallelism, backtracking, composition, and search-layer changes |
| Return the strongest target under best/ with an evidence summary | Internal candidate organization and the path used to reach the nominee |