MILO: Automated Harness Discovery via Orchestrated Multi-Agent Evolution
Organizations: Georgia Institute of Technology · AWS AI Labs · Carnegie Mellon University · Washington University in St. Louis
Abstract
Modern agentic systems combine an AI model with a harness that controls execution and environmental interactions. Harness design strongly affects long-horizon performance, yet its combinatorial search space demands substantial human effort that must be repeated as models change. Existing automated methods explore this space narrowly, optimizing only components such as prompts or skills or becoming trapped by fixed, exploitative search strategies. We introduce MILO (Meta-evolutionary Island Orchestration), a framework that co-evolves agent harnesses and the strategy used to discover them. MILO combines: (i) hierarchical lineage memory over island-based trees, using rejected mutations as negative evidence; (ii) per-island mutator agents that rewrite complete harnesses using global search history and parent-specific feedback; and (iii) an orchestrator that adapts search through lineage grafting and speciation, mutator reassignment and curriculum revision. Across Terminal-Bench 2.1, PaperBench, and DeepSWE, MILO-discovered harnesses outperform eight state-of-the-art harnesses and six search methods using frontier (Opus 4.8) and open-weight (gpt-oss-120b) models. With Opus 4.8, MILO improves resolution over its initial harness by , , and , respectively, compared with best prior-search gains of , , and . On Terminal-Bench 2.1, it achieves , exceeding the official leaderboard's top entry () while using 26% fewer tokens than its initial harness. On EinsteinArena open problems, MILO improves best-known upper bounds for Erdős minimum-overlap () and the first and third autocorrelation inequalities (; ).
Figures & tables
| Method | Level | Domain | Memory | Parent selection | Mutator |
| (0) Open-loop test-time scaling. No executed outcome feeds the next attempt, so gains cannot compound. feedback ✗ | |||||
| Best-of- ( Cobbe et al., 2021 ) | instance | Math | — | Best of i.i.d. draws | Single LLM |
| Chain-of-Thought ( Wei et al., 2022 ) | instance | Reasoning | — | — | Single LLM |
| Tree-of-Thoughts ( Yao et al., 2023a ) | instance | Planning | Transient thought tree | LLM-scored states | Single LLM |
| (I) Iterative refinement. Executed feedback steers one greedy lineage; no diversity-preserving population, so it stalls at local optima. feedback ✓ | |||||
| Reflexion ( Shinn et al., 2023 ) | instance | Code, QA | Verbal reflections | Latest attempt | Single LLM |
Appendix figures & tables2 assets
Supplementary material from the paper’s appendix.
Appendix
| Search reach | Failure mode addressed (mechanism) | |||||||||
| Method | Design | Pro- mpt | Skills/ mem | To- ols | Middle- ware | Premature completion | Deadline overrun | Lax verification | Missing know-how | RR@5 |
| reproduce-first rubric gate seed B12 | expert | ✓ | ✗ | ✗ | ✓ | ✓ rubric gate | ✗ | ✗ | ✗ | 74.1 |
| GEPA (prompt) | auto | ✓ | ✗ | ✗ | ✗ | ✓ rubric gate | ✗ | 78.6 | ||
| GEPA (optimize-anything) | auto | ✓ | ✗ | ✓ | ✓ | ✓ rubric gate | ✗ | 74.1 | ||
| A-Evolve (skill, memory) | auto | ✓ | ✓ | ✗ | ✗ | ✓ rubric gate | ✓ criteria skill | ✓ skills lib | 78.2 | |
| OpenEvolve | auto | ✓ | ✓ | ✓ | ✓ | ✓ rubric gate | ✗ | ✗ | ✗ | 76.1 |
| Problem (arena id) | Objective (minimization; lower is better) | AlphaEvolve (arena entry) | Prior best (Sept. 2026) | MILO | Margin | Min. impr. |
| Erdős minimum overlap (1) | minimize | 0.3809230351 | 0.3808585749 CodexProLong, 2026-08-15 | 0.3808567744 | ||
| First autocorrelation ineq. (2) | minimize , | 1.5052939684 | 1.5027436492 CodexProLong, 2026-08-14 | 1.5027435984 | ||
| Third autocorrelation ineq. (4) | minimize , signed | 1.4556427954 | 1.4508066395 Poolish, 2026-08-23 | 1.4488860143 |