Automated harness optimization can substantially improve LLM agents by iteratively updating their prompts, tool interfaces, and control logic from execution feedback. However, existing methods primarily optimize how the harness is updated while largely fixing which training scenarios generate the feedback that drives those updates. As the harness evolves, the scenarios most useful for further optimization can change, suggesting that the training curriculum itself should adapt alongside the harness. We formulate this missing dimension of harness optimization as an automated curriculum learning problem and introduce ActiveSaddler. ActiveSaddler models the evolving curriculum as a non-stationary bandit with dynamically instantiated optimization targets. It abstracts recurring failures into reusable failure-pattern arms, estimates the potential learning progress from further targeting each pattern, and adaptively balances revisiting known weaknesses with exploring unseen scenarios for new ones. Optimization outcomes continually update both the set of discovered failure patterns and their priorities, allowing the curriculum to co-evolve with the harness. Experiments on GAIA2 and Terminal-Bench 2.0 show that ActiveSaddler consistently discovers stronger harnesses, improving test Pass@1 by 4.4 and 7.5 percentage points over the same harness optimizer using a scenario order fixed before optimization, respectively. Ablations further show that these gains depend on dynamically constructing optimization targets, estimating their evolving utility, and balancing continued optimization with new failure discovery. Together, these results establish automated curriculum learning as a new crucial optimization dimension for harness optimization.
Figures & tables
Figure 1: Optimization efficiency across benchmarks. Best-so-far development accuracy versus cumulative task-agent rollouts for ActiveSaddler and baseline methods.
Figure 2: Overview of ActiveSaddler . At each iteration, the curriculum chooses between exploring unseen scenarios and revisiting an instantiated arm, optimizes the harness on the selected scenario batch, and updates its state based on the resulting execution and optimization outcomes.
Harness (Type)
Test-Split (300)
Default Agent (manual)
53.6 ± 1.1
GEPA
54.2 ± 2.2
Meta-Harness
54.2 ± 1.2
AutoSaddler
55.4 ± 1.2
AutoSaddler w/ Category Acc. Order
55.9 ± 1.3
AutoSaddler w/ Scenario Acc. Order
55.7 ± 1.2
Table 1: Test Pass@1 on GAIA2 and Terminal-Bench 2.0 (mean ± std. over three test-time executions). Parentheses show task counts; bold indicates the best result.
Figure 3: Effect of adaptive arm scoring. Adaptive scoring dynamically reallocates optimization priority and yields more effective repair of observed failures.
Figure 4: Adaptive exploration over the optimization trajectory. The policy chooses Draw when discovery is more valuable and Pull when promising repair opportunities remain.
Appendix figures & tables18 assets
Supplementary material from the paper’s appendix.
Appendix
Command
Purpose
When to use
Read operations
pattern list
Canonical arm table: score, observation history, scenario count, last observation, label
Start of session; orienting to the full arm population
Inspecting a specific arm before rating or patching
pattern score [--top-k <n>]
Top- k arms ranked by current score
Quick triage of the highest-priority arms
pattern scenarios --pattern-id <id>
Scenarios tagged to an arm, with co-occurring patterns
Checking an arm’s coverage and overlap
pattern history [--pattern-id <ids>] [--last-k <K>]
Complete per-arm pull history: patched attempts, all-pass skips, and failed attempts, with diagnoses, patch intents, dev-set impact, and per-scenario reflections
Before scoring or re-selecting an arm; assessing what prior repairs already attempted
Appendix
Table 2: Commands exposed by the pattern CLI, grouped into read operations for scoped views of the arm population and write operations for per-session registry updates. Each session receives only the subset of commands in its capability manifest.
Benchmark
Split Axis
Split
Task Group
# Tasks
GAIA2
Universe (persona)
Train
Universe 29
75
Dev
Universe 30
65
Test
Universe 21
107
Universe 22
112
Universe 27
81
Terminal-Bench 2.0
Random †
Train
—
30
Appendix
Table 3: Data splits across the two benchmarks. For GAIA2, we evaluate generalization under distribution shift by partitioning the benchmark such that train, development, and test sets contain tasks from disjoint Universes (personas), rather than using random task-level splits. Terminal-Bench 2.0 contains only 89 tasks across diverse domains and offers no natural grouping axis, so we adopt a uniform random partition.
Harness (Type)
GAIA2 Universe (Pass@1)
Avg.
21 (107)
22 (112)
27 (81)
Default Agent (manual)
54.5 ± 0.5
50.0 ± 1.5
57.2 ± 1.4
53.6 ± 1.1
AutoSaddler (Run 1)
57.0 ± 0.9
51.5 ± 2.2
58.8 ± 0.7
55.4 ± 1.2
AutoSaddler (Run 2)
57.6 ± 1.4
53.0 ± 1.0
60.1 ± 2.9
56.6 ± 1.0
AutoSaddler w/ Category Acc. Order (Run 1)
56.1 ± 3.4
53.9 ± 1.0
58.4 ± 1.9
55.9 ± 1.3
AutoSaddler w/ Category Acc. Order (Run 2)
57.0 ± 3.4
53.6 ± 0.9
58.4 ± 0.7
56.1 ± 1.0
Appendix
Table 4: Robustness to optimization stochasticity. Pass@1 on three held-out GAIA2 universes for harnesses obtained from two independent optimization runs of AutoSaddler, its difficulty-based fixed curricula, and ActiveSaddler. Performance on each universe is reported as mean ± standard deviation over three test-time executions, and the average is computed across all test scenarios.
Method
GAIA2 Universe (Pass@1)
Avg.
21 (107)
22 (112)
27 (81)
Default Agent (manual)
54.5 ± 0.5
50.0 ± 1.5
57.2 ± 1.4
53.6 ± 1.1
AutoSaddler
57.0 ± 0.9
51.5 ± 2.2
58.8 ± 0.7
55.4 ± 1.2
ActiveSaddler
60.4 ± 1.4
57.7 ± 2.6
61.7 ± 1.2
59.8 ± 1.0
ActiveSaddler w/ EMA Scoring
57.9 ± 2.5
53.0 ± 3.4
60.1 ± 0.7
56.7 ± 1.5
ActiveSaddler w/ UCB-AIR Controller
57.9 ± 3.4
50.9 ± 1.5
58.8 ± 3.1
55.6 ± 1.4
Appendix
Table 5: Comparison with alternative curriculum strategies. Pass@1 on three held-out GAIA2 universes when replacing either the Arm Prioritizer or Exploration Controller in ActiveSaddler with an alternative strategy while keeping the remaining curriculum components unchanged. Performance on each universe is reported as mean ± standard deviation over three test-time executions.
Method
GAIA2 Universe (Pass@1)
Avg.
21 (107)
22 (112)
27 (81)
Default Agent (manual)
54.5 ± 0.5
50.0 ± 1.5
57.2 ± 1.4
53.6 ± 1.1
GEPA
56.4 ± 1.4
50.0 ± 2.4
57.2 ± 3.6
54.2 ± 2.2
GEPA w/ ActiveSaddler
58.6 ± 0.5
54.2 ± 1.0
59.7 ± 0.7
57.2 ± 0.4
Appendix
Table 6: Applying ActiveSaddler to GEPA. Pass@1 on three held-out GAIA2 universes for the default agent, GEPA, and GEPA augmented with the ActiveSaddler curriculum. Performance on each universe is reported as mean ± standard deviation over three test-time executions. The average is computed across all test scenarios in the three universes.
Method
Generated
Rejected
Accepted
Time (s)
Cost ($)
LLM Calls
Output Tokens
Cache-Read Input Tokens
AutoSaddler
32
14
18
943
7.87
95.4
57,751
5,956,240
w/ Category Acc. Order
20
7
13
889
6.71
91.7
59,945
5,717,811
w/ Scenario Acc. Order
30
18
12
1,142
7.75
100.3
61,791
6,087,970
ActiveSaddler
37
18
19
1,581
11.44
142.2
89,525
8,704,858
Appendix
Table 7: Optimizer-side cost per generated patch. Runtime, monetary cost, LLM calls, and token usage are averaged per generated patch.
Method
LLM Calls
Input Tokens
Output Tokens
Time (s)
AutoSaddler
18.9
467,996
13,678
231.6
w/ Category Acc. Order
16.5
381,948
11,629
201.4
w/ Scenario Acc. Order
16.4
410,934
11,948
199.8
ActiveSaddler
17.5
410,960
12,340
212.0
Appendix
Table 8: Average task-agent evaluation cost per rollout.
Figure 5: End-to-end optimization efficiency. Best-so-far development accuracy against cumulative monetary cost, LLM calls, input tokens, and output tokens on GAIA2 (top) and Terminal-Bench 2.0 (bottom). Cumulative usage includes both optimizer-side and task-agent computation.
Figure 6: Case study on tnxtee (GAIA2), where a failure-pattern arm instantiated from the initial failure enables an unsuccessful repair to be retargeted and refined in a subsequent iteration.
Figure 7: Case study on 4bytkx (GAIA2), where retargeting the same failure-pattern arm corrects an executor hook that was initially inserted too late in the execution flow.
Figure 8: Failure hit rate by arm definition. Fraction of optimization iterations whose sampled mini-batch contains at least one scenario with an unresolved failure. Budgets are 1,400 rollouts for GAIA2 and 490 for TB2.
ID
Iter.
Induced Failure Pattern
Supporting Scenarios
1
1
No-attachment path defaults serialize as null instead of the oracle’s empty path string
8hgfug , aa11lh , tnxtee
2
1
Over-specific event notification wording diverges from concise oracle-style message content
8hgfug
3
3
Affirmative event-start email reply content is not canonicalized to the required acknowledgement structure
tnxtee
4
4
Multi-ambiguity user clarifications omit explicit questions for all unresolved decision dimensions
2qw4nk , qrxrry
5
4
Time-sensitive repeated actions drift against deadlines and stop before the required count is complete
71j6lf
6
4
Lookup tools reject raw identifiers that downstream action tools accept, causing resolution loops in deadline-bound tasks
71j6lf
Appendix
Table 9: Failure patterns induced by ActiveSaddler on GAIA2
ID
Iter.
Induced Failure Pattern
Supporting Scenarios
1
1
Premature completion on constrained file edits without rule-aware final validation
Security sanitizer tasks finalize after weak smoke tests without adversarial and clean-preservation validation
filter-js-from-html
4
4
Binary extraction solutions finalize after superficial output smoke tests without independent address-space convention validation
extract-elf
5
4
Agent self-validation leaves stale runtime artifacts that change verifier timing and hide required fresh-run output
make-doom-for-mips
6
4
Completion-only validation gates do not help long build/debug tasks that time out before attempting finalization
make-doom-for-mips
Appendix
Table 10: Failure patterns induced by ActiveSaddler on Terminal-Bench 2.0
Figure 9: Case study of a resolved weakness (P3 of GAIA2). After an accepted patch at Iteration 5, the arm’s priority drops sharply. A later pull finds all supporting scenarios successful and produces no further patch, after which the arm remains at low priority.
Figure 10: Case study of a recurring weakness (P8 of GAIA2). Early pulls produce no patch because the associated scenarios succeed when re-executed, yet the arm remains highly prioritized. The same failure pattern later appears in additional scenarios, providing new evidence that the weakness remains active; a subsequent pull produces an accepted patch. Cross markers indicate scenarios newly associated with P8.
Figure 11: Evolution of pull-probability mass across failure-pattern arms. The ten arms with the highest average pull probability over the run are shown individually, while the remaining 25 arms are aggregated into Other arms . The upper panel shows the number of instantiated arms, which grows from 6 to 35 as new failure patterns are discovered. Pull probabilities are shown only at the 34 iterations where arm selection is performed.
Figure 12: Pull-probability trajectories for all failure-pattern arms. Rows show all 35 discovered arms ordered by instantiation time, and columns correspond to the 34 iterations in which arm selection is performed. Cell intensity denotes pull probability, orange markers indicate the arm sampled at each iteration, and gray cells denote iterations before an arm entered the pool.
Figure 13: Dependence of exploration decisions on unresolved weaknesses. Ratio of the average number of unresolved arms at Pull to Draw decisions. Values above one indicate more unresolved weaknesses at Pull , while values near one indicate little difference between the two decision types.
Harness design plays a critical role in agent performance by shaping how large language models (LLMs) perceive, reason over, and act within executable environments. Recent work has proposed automatic harness evolution, which iteratively improves the harness from agent--environment interactions. However, existing methods often overfit to the evolution tasks, rely exclusively on trajectory-derived signals, and optimize harness components jointly, causing interference across components. We propose HarnessCompass, a novel automatic harness evolution framework built around constrained evolution, proactive feedback, and component-wise optimization. HarnessCompass first enforces global constraints on evolution, restricting modifications to task-agnostic harness changes that generalize beyond the evolution tasks. It then augments trajectory-derived evidence with proactive first-person feedback from the agent about harness usage, yielding richer signals for evolution. Finally, it decouples the optimization of different harness components before consolidating them into a unified harness, reducing cross-component interference while preserving component synergy. On SWE-bench Verified with GPT-5.4, HarnessCompass improves Pass@1 from 54% to 66% in only 5 evolution iterations, outperforming AHE in both effectiveness and evolution efficiency. In addition, the evolved harness transfers effectively to held-out tasks and other models, demonstrating substantially stronger generalization than prior automatic harness evolution methods.
Luan Zhang, Ruochen Zhou, Dandan Song +9
Beijing Institute of Technology, China · City University of Hong Kong, China · Independent, China
As LLMs are increasingly deployed within agentic systems, their capabilities depend not only on the model weights but also on the harness: the prompts, tools, control flow, memory, and orchestration code surrounding them. This makes automated harness optimization -- the iterative and evaluation-guided improvement of a harness by an AI system -- both an important route to improving AI systems and a demanding capability for AI systems themselves. Yet the community lacks a common protocol for measuring how well frontier LLMs perform at this task. We introduce HarnessOpt-Bench, a benchmark for end-to-end harness optimization under expensive and stochastic evaluation. An optimizer, an LLM paired with a coding harness, receives a target agent's seed harness, graded evaluation feedback, and a fixed target-evaluation budget. It edits the harness and nominates a final candidate, which is scored by its normalized gain over the seed on a held-out test partition that remains inaccessible throughout search. A trusted execution environment enforces the evaluation boundary, meters target-agent resource use, and preserves candidate versions for audit. We evaluate 5 frontier LLMs as optimizers both under a shared coding harness and under their native harnesses across 4 downstream tasks, over 111 scored runs. Experiment results show that optimizer models separate more than the coding harnesses they act through, native harnesses are not consistently superior, and gains vary substantially across tasks and seed regimes. These results establish harness optimization as a measurable and discriminative capability with large space for improvement.
Post-training agents for automated AI research requires optimizing not only model parameters, but also the runtime harness that shapes how research trajectories are generated, evaluated, and learned from. Existing pipelines typically train models under a fixed harness, including prompts, tools, skills, middleware, and memory, while leaving the data-generating process outside the optimization objective. This creates a mismatch between model updates and the static scaffolding that determines trajectory quality. We introduce Co-Harness, a framework that jointly optimizes the agent harness and model parameters during post-training. Co-Harness alternates between harness optimization and model optimization. An LLM-based HarnessCritic analyzes failed trajectories, identifies harness-level failure modes, and proposes validated local updates. The model is then fine-tuned on high-quality trajectories generated by the improved harness, distilling effective scaffolding into model parameters. A 200+ hour autonomous case study further shows that Co-Harness can recover from system crashes, improve inference efficiency, and discover ensemble strategies without human intervention. These results suggest that joint harness and model optimization is an effective way to improve agents beyond fixed-harness post-training.