Adaptive-GEPA: Make Your Harness Fit Heterogeneous Requests
Organizations: University of Texas at Austin · University of California, Los Angeles · University of California, San Diego · Carnegie Mellon University · Shenzhen Institutes of Advanced Technology · Microsoft
Abstract
Reflective optimizers such as GEPA improve language model prompts from execution traces and evaluator feedback; full-program extensions can also rewrite tools and control flow. In practice, a user hands the same endpoint heterogeneous requests whose effective solutions require different tools, reasoning modes, and control flow. Optimizing one shared program leaves this division of work implicit in source-code search, while optimizing a separate program per request family fixes it beforehand. We introduce Adaptive-GEPA, which learns both how to divide requests and how to solve them. It evolves a router and a library of specialist programs under one search budget. The router's instructions, each specialist's description, and its program code are plain, human-readable text, edited from feedback. To combine branches, it aligns specialists by the requests they handle and inherits descriptions together with programs. On a fixed mixture of four task families, the reported Qwen3-8B run evolves four experts without supplying family labels to the router or reflection model; its routing matches the task partition on all 651 test requests. Its family-mean test score (x100) rises from 52.6 to 70.6, compared with 62.5 for GEPA's full-program adapter and 54.0 for GRPO at a nominal budget of 18,000 scored calls. These counts do not equate total compute. Figure 1 summarizes the learning curves, final test scores, and routing agreement.
Figures & tables
| Method | Search variables | Request allocation | Update organization |
|---|---|---|---|
| Original Gepa ( Agrawal et al., 2026 ) | Module prompts | Supplied workflow; prompts editable | Round-robin edits; optional module merges |
| MoP ( Wang et al., 2024 ) | Cluster count, instructions, demo assignments | Nearest embedding centroid | Clustering, then prompt assignment |
| Gepa-Fpa -ALL | One complete program | Dispatch can be generated in source | Full-program edits |
| Per-family Gepa-Fpa | A program per known family | True task labels | Separate searches |
| Adaptive- Gepa | Router, descriptions, programs, active slots | Router reads requests and descriptions | Feedback selects action and target; aligned merges |
| Budget | Test score ( ) | Mean ( ) | |||||
|---|---|---|---|---|---|---|---|
| Method | (calls) | HotpotQA | AIME | PUPA | LB-IF | Family | Example |
| No task labels supplied at inference | |||||||
| Generic seed | — | 26.0 | 32.0 | 76.7 | 75.8 | 52.6 | 51.1 |
| MoP | 1k | 31.2 | 16.0 | 78.5 | 79.2 | 51.2 | 54.0 |
| GRPO | 18k | 28.5 | 33.3 | 76.3 | 78.1 | 54.0 | 52.5 |
| GRPO ( budget) | 90k | 26.2 | 39.3 | 78.9 | 80.6 | 56.3 | 53.0 |
| Budget | Test score ( ) | Mean ( ) | |||||
|---|---|---|---|---|---|---|---|
| Method | (calls) | HotpotQA | AIME | PUPA | LB-IF | Family | Example |
| No task labels supplied at inference | |||||||
| Generic seed | — | 53.3 | 30.7 | 80.5 | 75.3 | 59.9 | 64.9 |
| MoP | 1k | 47.3 | 27.3 | 75.4 | 73.3 | 55.8 | 59.9 |
| Gepa-Fpa -ALL | 18k | 58.1 | 36.0 | 93.0 | 92.3 | 69.8 | 74.2 |
| Adaptive- Gepa | 18k | 62.5 | 40.0 | 90.8 | 92.6 | 71.5 | 75.7 |
| Family | MoP | Ours |
|---|---|---|
| HotpotQA | 94.7 | 100.0 |
| AIME | 100.0 | 100.0 |
| PUPA | 88.2 | 100.0 |
| LB-IF | 97.0 | 100.0 |
| Overall | 93.1 | 100.0 |
| Local edit | Proposed | Accepted |
|---|---|---|
| Activate specialist | 22 | 16 |
| Rewrite program | 38 | 27 |
| Rewrite router | 12 | 8 |
| Rewrite description | 4 | 1 |
| Total | 76 | 52 |
Appendix figures & tables12 assets
Supplementary material from the paper’s appendix.
Appendix
| Operation | Traces | Router | All descriptors | Target source | Sibling sources |
|---|---|---|---|---|---|
| Action selection | ✓ | ✓ | ✓ | – | – |
| Activate slot | ✓ | ✓ | ✓ | – | – |
| Rewrite program | ✓ | ✓ | ✓ | ✓ | – |
| Rewrite router | ✓ | ✓ | ✓ | – | – |
| Rewrite descriptor | ✓ | ✓ | ✓ | ✓ | – |
| Quarantine slot | – | – | – | – | – |
| Component | Gepa-Fpa -ALL | Adaptive- Gepa |
|---|---|---|
| Dispatch | Python branches for privacy, writing, mathematics, retrieval; generic fallback | LM router reads descriptions and selects an active program or fixed fallback |
| Retrieval | Heuristic and LM queries; ColBERT; code-generated bridge queries; grounded answer | Heuristic and LM queries; ColBERT; evidence-conditioned follow-up; draft and LM evidence check |
| Mathematics | Deterministic solvers; otherwise one reasoning attempt and answer normalization | Deterministic solvers; otherwise four attempts, answer aggregation, and LM review |
| Privacy | Regex replacement of email, phone, and a specific disability phrase | LM sanitization and review; identifier cleanup; nonempty fallback request |
| Writing | LM draft; deterministic keyword, length, structure, and wrapper repair | Constraint checklist; LM draft; deterministic repairs and word-limit loop; fixed-choice bypass |
| Profile | One global deep profile | Deep for mathematics and writing; fast for retrieval, privacy, and routing |
| Qwen3-8B | GPT-4.1 | |||||
|---|---|---|---|---|---|---|
| Excluded family | FPA | Ours | FPA | Ours | ||
| None | 62.47 | 70.58 | +8.11 | 69.84 | 71.46 | +1.62 |
| HotpotQA | 65.27 | 74.71 | +9.44 | 73.76 | 74.44 | +0.68 |
| AIME | 71.07 | 78.32 | +7.26 | 81.12 | 81.95 | +0.83 |
| PUPA | 60.69 | 63.50 | +2.81 | 62.12 | 65.03 | +2.91 |
| LiveBench-IF | 52.83 | 65.76 | +12.93 | 62.36 | 64.43 | +2.07 |
| Test score | Test mean | ||||||
|---|---|---|---|---|---|---|---|
| Checkpoint | Val. mean | HotpotQA | AIME | PUPA | LB-IF | Family | Example |
| Seed | 50.4 | 26.6 | 36.7 | 78.5 | 74.3 | 54.0 | 52.0 |
| 115 | 60.2 | 28.5 | 36.0 | 80.2 | 81.9 | 56.6 | 54.6 |
| 180 (selected) | 62.8 | 26.2 | 39.3 | 78.9 | 80.6 | 56.3 | 53.0 |
| 235 (final) | 59.7 | 27.9 | 40.7 | 76.8 | 84.1 | 57.4 | 53.7 |
| Archived child | Count | Higher | Lower | Median | Range of |
|---|---|---|---|---|---|
| Active pairs from both parents | 10 | 10 | 0 | +3.20 | |
| Identical to a parent | 13 | 7 | 6 | +0.08 | |
| Only inactive code differs | 3 | 1 | 2 | ||
| All accepted merges | 26 | 18 | 8 | +0.67 |
| Seed | Program | HotpotQA | AIME | LB-IF | Mean | |
|---|---|---|---|---|---|---|
| 42 | Initial | 28.7 | 31.3 | 67.6 | 42.56 | — |
| 42 | Selected | 60.0 | 40.0 | 71.5 | 57.16 | +14.60 |
| 43 | Initial | 25.4 | 34.0 | 70.3 | 43.24 | — |
| 43 | Selected | 62.6 | 41.3 | 80.2 | 61.38 | +18.14 |
| HotpotQA | AIME | PUPA | LB-IF | Mean | |
|---|---|---|---|---|---|
| Parent 52 | 62.8 | 35.6 | 86.5 | 89.9 | 68.7 |
| Parent 66 | 73.6 | 33.3 | 88.3 | 73.7 | 67.2 |
| Index child (measured) | 38.3 | 35.6 | 88.1 | 89.9 | 63.0 |
| Aligned child (projected) | 73.6 | 35.6 | 86.5 | 89.9 | 71.4 |
| Method | Editable unit | Selection or inheritance |
|---|---|---|
| AutoPDL ( Spiess et al., 2025 ) | Prompting patterns, instructions, and demonstrations assembled into executable PDL programs. | Successive halving selects configurations within a predefined search space. |
| JTPRO ( Ghoshal et al., 2026 ) | Global instructions and tool schemas, including argument descriptions. | Merges edits with the best context while maintaining correspondence through tool identity. |
| Adaptive Auto-Harness ( Liu et al., 2026b ) | Harness branches with prompts, skills, tools, memory, and infrastructure for an open-ended task stream. | Git branches preserve alternative harnesses and their histories; a router selects a branch for a request. |
| Adaptive- Gepa | A router and active description–program pairs, jointly evolved on a fixed request mixture. | Per-request selection retains alternatives; footprint alignment supplies correspondence for pair inheritance across candidates. |