Organizations: University of Texas at Austin · University of California, Los Angeles · University of California, San Diego · Carnegie Mellon University · Shenzhen Institutes of Advanced Technology · Microsoft
Reflective optimizers such as GEPA improve language model prompts from execution traces and evaluator feedback; full-program extensions can also rewrite tools and control flow. In practice, a user hands the same endpoint heterogeneous requests whose effective solutions require different tools, reasoning modes, and control flow. Optimizing one shared program leaves this division of work implicit in source-code search, while optimizing a separate program per request family fixes it beforehand. We introduce Adaptive-GEPA, which learns both how to divide requests and how to solve them. It evolves a router and a library of specialist programs under one search budget. The router's instructions, each specialist's description, and its program code are plain, human-readable text, edited from feedback. To combine branches, it aligns specialists by the requests they handle and inherits descriptions together with programs. On a fixed mixture of four task families, the reported Qwen3-8B run evolves four experts without supplying family labels to the router or reflection model; its routing matches the task partition on all 651 test requests. Its family-mean test score (x100) rises from 52.6 to 70.6, compared with 62.5 for GEPA's full-program adapter and 54.0 for GRPO at a nominal budget of 18,000 scored calls. These counts do not equate total compute. Figure 1 summarizes the learning curves, final test scores, and routing agreement.
Figures & tables
Figure 1: One routed search improves a mixed workload. (a) Best validation means; stars mark test scores. (b) Per-task test scores. (c) Post-hoc family agreement ( 85 – 100% axis). Nominal budgets are 18,000 scored calls for search and GRPO, and 1,010 for MoP.
Figure 2: Adaptive- Gepa learns both the division of work and the programs that carry it out. The actor LLM (blue) routes requests and supplies model calls inside programs. The reflection LLM (purple) selects actions and generates any needed code or text edits. The search algorithm matches experts by request overlap, resolves conflicts using scores on shared requests, and retains each role–program pair intact. The schematic score bars and check illustrate expert selection. Passing proposals receive full validation and enter the archive. The snapshots show growth and edits across iterations; gray cards denote the fixed fallback, plus signs additions, and pencils edits. No task labels are supplied to the router or reflection model .
Method
Search variables
Request allocation
Update organization
Original Gepa ( Agrawal et al., 2026 )
Module prompts
Supplied workflow; prompts editable
Round-robin edits; optional module merges
MoP ( Wang et al., 2024 )
Cluster count, instructions, demo assignments
Nearest embedding centroid
Clustering, then prompt assignment
Gepa-Fpa -ALL
One complete program
Dispatch can be generated in source
Full-program edits
Per-family Gepa-Fpa
A program per known family
True task labels
Separate searches
Adaptive- Gepa
Router, descriptions, programs, active slots
Router reads requests and descriptions
Feedback selects action and target; aligned merges
Table 1: Adaptive- Gepa makes routing, responsibilities, and programs separate search targets. FPA denotes GEPA’s full-program adapter; MoP is Mixture-of-Prompts. Original GEPA is a design reference; the other rows describe the evaluated configurations.
Budget
Test score ( ×100 ) ↑
Mean ( ×100 ) ↑
Method
(calls)
HotpotQA n=300
AIME n=30
PUPA n=221
LB-IF n=100
Family
Example
No task labels supplied at inference
Generic seed
—
26.0
32.0
76.7
75.8
52.6
51.1
MoP
≈ 1k
31.2
16.0
78.5
79.2
51.2
54.0
GRPO
18k
28.5
33.3
76.3
78.1
54.0
52.5
GRPO ( 5× budget)
90k
26.2
39.3
78.9
80.6
56.3
53.0
Table 2: One joint search approaches independently optimized, label-routed programs. Qwen3-8B test scores (seed 43). Means weight families (primary) or examples equally; n is the test-set size. Bold marks the best score without task labels. Budgets count scored calls; program search uses nominal caps, and the 90k GRPO cutoff is retrospective.
Budget
Test score ( ×100 ) ↑
Mean ( ×100 ) ↑
Method
(calls)
HotpotQA n=300
AIME n=30
PUPA n=221
LB-IF n=100
Family
Example
No task labels supplied at inference
Generic seed
—
53.3
30.7
80.5
75.3
59.9
64.9
MoP
≈ 1k
47.3
27.3
75.4
73.3
55.8
59.9
Gepa-Fpa -ALL
18k
58.1
36.0
93.0
92.3
69.8
74.2
Adaptive- Gepa
18k
62.5
40.0
90.8
92.6
71.5
75.7
Table 3: GPT-4.1 narrows the advantage over single-program evolution. Conventions follow Table 2 . Resumptions and a judge outage limit the evidence (Appendix L ).
Figure 3: The final harness learns different procedures. Qwen3-8B candidate 74 pairs descriptions (left) and programs (right), with post-hoc task labels below. Slot 5 is inactive; fallback is fixed.
Family
MoP
Ours
HotpotQA
94.7
100.0
AIME
100.0
100.0
PUPA
88.2
100.0
LB-IF
97.0
100.0
Overall
93.1
100.0
Table 4: Routing matches the task partition. Family agreement (%) under a post-hoc majority-family mapping.
Figure 4: Evolving libraries. Nodes show validation means; arrows link parents.
Figure 5: Slot use. 45 requests per task; 180 for fallback. Slots are mapped to tasks after search.
Local edit
Proposed
Accepted
Activate specialist
22
16
Rewrite program
38
27
Rewrite router
12
8
Rewrite description
4
1
Total
76
52
Table 5: Local edit counts. All branches of the Qwen3-8B run; merges excluded.
Appendix figures & tables12 assets
Supplementary material from the paper’s appendix.
Appendix
Operation
Traces
Router
All descriptors
Target source
Sibling sources
Action selection
✓
✓
✓
–
–
Activate slot
✓
✓
✓
–
–
Rewrite program
✓
✓
✓
✓
–
Rewrite router
✓
✓
✓
–
–
Rewrite descriptor
✓
✓
✓
✓
–
Quarantine slot
–
–
–
–
–
Appendix
Table 6: Information visible to each reflection operation.
Component
Gepa-Fpa -ALL
Adaptive- Gepa
Dispatch
Python branches for privacy, writing, mathematics, retrieval; generic fallback
LM router reads descriptions and selects an active program or fixed fallback
Retrieval
Heuristic and LM queries; ColBERT; code-generated bridge queries; grounded answer
Heuristic and LM queries; ColBERT; evidence-conditioned follow-up; draft and LM evidence check
Mathematics
Deterministic solvers; otherwise one reasoning attempt and answer normalization
Deterministic solvers; otherwise four attempts, answer aggregation, and LM review
Privacy
Regex replacement of email, phone, and a specific disability phrase
LM sanitization and review; identifier cleanup; nonempty fallback request
Writing
LM draft; deterministic keyword, length, structure, and wrapper repair
Deep for mathematics and writing; fast for retrieval, privacy, and routing
Appendix
Table 7: Both searches learn multi-stage procedures. Static comparison of the validation-selected Qwen programs: Gepa-Fpa -ALL candidate 21 and Adaptive- Gepa candidate 74. Conditional paths are condensed; shared evaluation-side models are omitted.
Figure 6: Edits change program use. Top: recorded executions. Bottom: before/after description and router excerpts; ellipses mark omissions.
Qwen3-8B
GPT-4.1
Excluded family
FPA
Ours
Δ
FPA
Ours
Δ
None
62.47
70.58
+8.11
69.84
71.46
+1.62
HotpotQA
65.27
74.71
+9.44
73.76
74.44
+0.68
AIME
71.07
78.32
+7.26
81.12
81.95
+0.83
PUPA
60.69
63.50
+2.81
62.12
65.03
+2.91
LiveBench-IF
52.83
65.76
+12.93
62.36
64.43
+2.07
Appendix
Table 8: The mean advantage remains after excluding any one family. Scores ( ×100 ) use fixed final candidates and equally weight the remaining families. Δ is Ours minus FPA.
Test score
Test mean
Checkpoint
Val. mean
HotpotQA
AIME
PUPA
LB-IF
Family
Example
Seed
50.4
26.6
36.7
78.5
74.3
54.0
52.0
115
60.2
28.5
36.0
80.2
81.9
56.6
54.6
180 (selected)
62.8
26.2
39.3
78.9
80.6
56.3
53.0
235 (final)
59.7
27.9
40.7
76.8
84.1
57.4
53.7
Appendix
Table 9: Validation and test scores favor different checkpoints. All scores are multiplied by 100. Both the 90k cutoff and full-run validation select step 180. Steps 115 and 235 are post-hoc diagnostics; step 235 lies beyond 90k. Means follow Table 2 .
Archived child
Count
Higher
Lower
Median Δ
Range of Δ
Active pairs from both parents
10
10
0
+3.20
[+0.34,+6.54]
Identical to a parent
13
7
6
+0.08
[−2.05,+2.55]
Only inactive code differs
3
1
2
−2.85
[−3.05,+0.49]
All accepted merges
26
18
8
+0.67
[−3.05,+6.54]
Appendix
Table 10: Ten accepted merges combine distinct active pairs from both parents. All 26 accepted merges in the main run. Δ is the recorded full-validation mean minus the stronger parent’s mean, multiplied by 100. Classes describe the archived source after cleanup; scores precede cleanup.
Figure 7: Observed request overlap aligns specialists across different slot layouts. Left: Jaccard overlap of the two parents’ recorded validation footprints. Right: measured family scores for both parents and their merged child. Large off-diagonal entries show that slot position alone does not identify a responsibility. Task labels only annotate the plot; Instr. means instruction following.
Figure 8: Merging aligns specialists, selects components, and evaluates a new library. Archived request overlap aligns parents 55 and 59; shared-request scores select description–program pairs when both changed. Validation precedes archiving and frontier refresh. Colors show post-hoc roles.
Seed
Program
HotpotQA
AIME
LB-IF
Mean
Δ
42
Initial
28.7
31.3
67.6
42.56
—
42
Selected
60.0
40.0
71.5
57.16
+14.60
43
Initial
25.4
34.0
70.3
43.24
—
43
Selected
62.6
41.3
80.2
61.38
+18.14
Appendix
Table 11: Historical runs improve over their own initial programs. Results for predecessor v4, not repeated runs of the final method. Scores are multiplied by 100; the mean excludes PUPA. Δ compares each selected program with its own initial program.
HotpotQA
AIME
PUPA
LB-IF
Mean
Parent 52
62.8
35.6
86.5
89.9
68.7
Parent 66
73.6
33.3
88.3
73.7
67.2
Index child (measured)
38.3
35.6
88.1
89.9
63.0
Aligned child (projected)
73.6
35.6
86.5
89.9
71.4
Appendix
Table 12: Index-based merging can lose a retrieval specialist. Scores are multiplied by 100. Parent and index-child scores are measured; the aligned-child row is projected from retained parent programs. Candidate indices refer to a historical run, not the main run.
Figure 9: Search evolves entire candidate libraries. All 79 records and 104 parent edges are shown. Each card contains a router above five description/program pairs; gray pairs are inactive but stored. The fixed fallback is retained but omitted. Teal edges trace candidate 74’s ancestry; gold borders mark recorded final-frontier membership. Candidate 74 has a separate selected border; its frontier status is unrecorded. Labels give IDs and validation means. Candidates 73 and 74 contain identical components despite their different scores.
Method
Editable unit
Selection or inheritance
AutoPDL ( Spiess et al., 2025 )
Prompting patterns, instructions, and demonstrations assembled into executable PDL programs.
Successive halving selects configurations within a predefined search space.
JTPRO ( Ghoshal et al., 2026 )
Global instructions and tool schemas, including argument descriptions.
Merges edits with the best context while maintaining correspondence through tool identity.
Adaptive Auto-Harness ( Liu et al., 2026b )
Harness branches with prompts, skills, tools, memory, and infrastructure for an open-ended task stream.
Git branches preserve alternative harnesses and their histories; a router selects a branch for a request.
Adaptive- Gepa
A router and active description–program pairs, jointly evolved on a fixed request mixture.
Per-request selection retains alternatives; footprint alignment supplies correspondence for pair inheritance across candidates.
Appendix
Table 13: Related methods expose different units for optimization. Each search object determines the editable components and how alternatives are selected or combined.
Prompt optimizers automate the search for prompts that improve language-model performance, but existing methods rely on a predefined optimization procedure: the algorithm determines which candidates to explore and how the search progresses, while the language model generates or refines prompt proposals. We introduce RLMOpt, a prompt optimizer that makes the search policy itself language-model-driven through a recursive language model (RLM). The RLM agent operates over a tool-based environment, inspecting task information, analyzing failures, generating candidates, allocating evaluation budget, and deciding when to stop. A deterministic harness complements the agent by enforcing objective scoring, Pareto-based selection, and regression constraints. We evaluate RLMOpt across four benchmarks spanning structured clinical information extraction (Chia), multi-hop question answering (HotpotQA), verifiable instruction following (IFBench-2025), and multi-turn tool-calling agents (BFCL). In a matched comparison at a single seed, RLMOpt obtains the best held-out score on all four benchmarks and leads the four-task mean (0.610 against 0.589 for GEPA). Repeating each benchmark across seeds yields 11 matched benchmark-seed comparisons, in which RLMOpt outperforms GEPA in 9 cases. Across all 11 runs, it never produced a prompt that underperformed its seed, whereas GEPA fell below its starting point twice. It is also more efficient, achieving these results with fewer search rollouts while producing prompts that are 27-79% the size of those produced by GEPA. Our results further show that optimization gains are determined primarily by the headroom available in the seed prompt, rather than by the search budget. Efficient optimization therefore depends on reaching the available headroom reliably and with minimal search
A language-model agent is jointly defined by its model and its harness, the executable program that organizes model calls, tool use, and information flow. Because different tasks call for different ways of organizing these operations, the harness needs to be adapted using feedback from the task at hand. We introduce harness learning, which trains a proposer model to revise a solver's harness using execution feedback. We formulate this process as meta-learning over executable programs, with harness revisions playing the role of weight updates in gradient-based adaptation. We train the proposer with reinforcement learning, using the task performance of revised harnesses as the reward. At test time, the proposer uses feedback from successive executions on a new task to refine the harness, without performing any parameter-space update. Experiments on reasoning and multi-hop question answering show that harness learning improves revision quality and that the ability to adapt at test time transfers to unseen tasks. Policies trained on individual revisions can continue improving harnesses over multiple rounds, while the benefits of training on revision sequences vary across settings. These findings suggest a path towards continually learning agents that turn accumulated experience into generalizable improvements.
Can a single LLM-based optimization system match specialized tools across fundamentally different domains? We show that when optimization problems are formulated as improving a text artifact evaluated by a scoring function, a single AI-based optimization system-supporting single-task search, multi-task search with cross-problem transfer, and generalization to unseen inputs-achieves state-of-the-art results across six diverse tasks. Our system discovers agent architectures that nearly triple Gemini Flash's ARC-AGI accuracy (32.5% to 89.5%), finds scheduling algorithms that cut cloud costs by 40%, generates CUDA kernels where 87% match or beat PyTorch, and outperforms AlphaEvolve's reported circle packing solution (n=26). Ablations across three domains reveal that actionable side information yields faster convergence and substantially higher final scores than score-only feedback, and that multi-task search outperforms independent optimization given equivalent per-problem budget through cross-task transfer, with benefits scaling with the number of related tasks. Together, we show for the first time that text optimization with LLM-based search is a general-purpose problem-solving paradigm, unifying tasks traditionally requiring domain-specific algorithms under a single framework. We open-source optimize_anything with support for multiple backends as part of the GEPA project at https://github.com/gepa-ai/gepa .