ActiveSaddler: Automated Curriculum Learning for Agent Harness Optimization
Organizations: POSTECH · KAIST · Microsoft
Abstract
Automated harness optimization can substantially improve LLM agents by iteratively updating their prompts, tool interfaces, and control logic from execution feedback. However, existing methods primarily optimize how the harness is updated while largely fixing which training scenarios generate the feedback that drives those updates. As the harness evolves, the scenarios most useful for further optimization can change, suggesting that the training curriculum itself should adapt alongside the harness. We formulate this missing dimension of harness optimization as an automated curriculum learning problem and introduce ActiveSaddler. ActiveSaddler models the evolving curriculum as a non-stationary bandit with dynamically instantiated optimization targets. It abstracts recurring failures into reusable failure-pattern arms, estimates the potential learning progress from further targeting each pattern, and adaptively balances revisiting known weaknesses with exploring unseen scenarios for new ones. Optimization outcomes continually update both the set of discovered failure patterns and their priorities, allowing the curriculum to co-evolve with the harness. Experiments on GAIA2 and Terminal-Bench 2.0 show that ActiveSaddler consistently discovers stronger harnesses, improving test Pass@1 by 4.4 and 7.5 percentage points over the same harness optimizer using a scenario order fixed before optimization, respectively. Ablations further show that these gains depend on dynamically constructing optimization targets, estimating their evolving utility, and balancing continued optimization with new failure discovery. Together, these results establish automated curriculum learning as a new crucial optimization dimension for harness optimization.
Figures & tables
| Harness (Type) | Test-Split (300) |
| Default Agent (manual) | 53.6 ± 1.1 |
| GEPA | 54.2 ± 2.2 |
| Meta-Harness | 54.2 ± 1.2 |
| AutoSaddler | 55.4 ± 1.2 |
| AutoSaddler w/ Category Acc. Order | 55.9 ± 1.3 |
| AutoSaddler w/ Scenario Acc. Order | 55.7 ± 1.2 |
Appendix figures & tables18 assets
Supplementary material from the paper’s appendix.
Appendix
| Command | Purpose | When to use |
| Read operations | ||
| pattern list | Canonical arm table: score, observation history, scenario count, last observation, label | Start of session; orienting to the full arm population |
| pattern show <id> | Pattern details: label, score, tagged tuples, per-scenario root causes, observations | Inspecting a specific arm before rating or patching |
| pattern score [--top-k <n>] | Top- arms ranked by current score | Quick triage of the highest-priority arms |
| pattern scenarios --pattern-id <id> | Scenarios tagged to an arm, with co-occurring patterns | Checking an arm’s coverage and overlap |
| pattern history [--pattern-id <ids>] [--last-k <K>] | Complete per-arm pull history: patched attempts, all-pass skips, and failed attempts, with diagnoses, patch intents, dev-set impact, and per-scenario reflections | Before scoring or re-selecting an arm; assessing what prior repairs already attempted |
| Benchmark | Split Axis | Split | Task Group | # Tasks |
| GAIA2 | Universe (persona) | Train | Universe 29 | 75 |
| Dev | Universe 30 | 65 | ||
| Test | Universe 21 | 107 | ||
| Universe 22 | 112 | |||
| Universe 27 | 81 | |||
| Terminal-Bench 2.0 | Random † | Train | — | 30 |
| Harness (Type) | GAIA2 Universe (Pass@1) | Avg. | ||
| 21 (107) | 22 (112) | 27 (81) | ||
| Default Agent (manual) | 54.5 ± 0.5 | 50.0 ± 1.5 | 57.2 ± 1.4 | 53.6 ± 1.1 |
| AutoSaddler (Run 1) | 57.0 ± 0.9 | 51.5 ± 2.2 | 58.8 ± 0.7 | 55.4 ± 1.2 |
| AutoSaddler (Run 2) | 57.6 ± 1.4 | 53.0 ± 1.0 | 60.1 ± 2.9 | 56.6 ± 1.0 |
| AutoSaddler w/ Category Acc. Order (Run 1) | 56.1 ± 3.4 | 53.9 ± 1.0 | 58.4 ± 1.9 | 55.9 ± 1.3 |
| AutoSaddler w/ Category Acc. Order (Run 2) | 57.0 ± 3.4 | 53.6 ± 0.9 | 58.4 ± 0.7 | 56.1 ± 1.0 |
| Method | GAIA2 Universe (Pass@1) | Avg. | ||
| 21 (107) | 22 (112) | 27 (81) | ||
| Default Agent (manual) | 54.5 ± 0.5 | 50.0 ± 1.5 | 57.2 ± 1.4 | 53.6 ± 1.1 |
| AutoSaddler | 57.0 ± 0.9 | 51.5 ± 2.2 | 58.8 ± 0.7 | 55.4 ± 1.2 |
| ActiveSaddler | 60.4 ± 1.4 | 57.7 ± 2.6 | 61.7 ± 1.2 | 59.8 ± 1.0 |
| ActiveSaddler w/ EMA Scoring | 57.9 ± 2.5 | 53.0 ± 3.4 | 60.1 ± 0.7 | 56.7 ± 1.5 |
| ActiveSaddler w/ UCB-AIR Controller | 57.9 ± 3.4 | 50.9 ± 1.5 | 58.8 ± 3.1 | 55.6 ± 1.4 |
| Method | GAIA2 Universe (Pass@1) | Avg. | ||
| 21 (107) | 22 (112) | 27 (81) | ||
| Default Agent (manual) | 54.5 ± 0.5 | 50.0 ± 1.5 | 57.2 ± 1.4 | 53.6 ± 1.1 |
| GEPA | 56.4 ± 1.4 | 50.0 ± 2.4 | 57.2 ± 3.6 | 54.2 ± 2.2 |
| GEPA w/ ActiveSaddler | 58.6 ± 0.5 | 54.2 ± 1.0 | 59.7 ± 0.7 | 57.2 ± 0.4 |
| Method | Generated | Rejected | Accepted | Time (s) | Cost ($) | LLM Calls | Output Tokens | Cache-Read Input Tokens |
| AutoSaddler | 32 | 14 | 18 | 943 | 7.87 | 95.4 | 57,751 | 5,956,240 |
| w/ Category Acc. Order | 20 | 7 | 13 | 889 | 6.71 | 91.7 | 59,945 | 5,717,811 |
| w/ Scenario Acc. Order | 30 | 18 | 12 | 1,142 | 7.75 | 100.3 | 61,791 | 6,087,970 |
| ActiveSaddler | 37 | 18 | 19 | 1,581 | 11.44 | 142.2 | 89,525 | 8,704,858 |
| Method | LLM Calls | Input Tokens | Output Tokens | Time (s) |
| AutoSaddler | 18.9 | 467,996 | 13,678 | 231.6 |
| w/ Category Acc. Order | 16.5 | 381,948 | 11,629 | 201.4 |
| w/ Scenario Acc. Order | 16.4 | 410,934 | 11,948 | 199.8 |
| ActiveSaddler | 17.5 | 410,960 | 12,340 | 212.0 |
| ID | Iter. | Induced Failure Pattern | Supporting Scenarios |
| 1 | 1 | No-attachment path defaults serialize as null instead of the oracle’s empty path string | 8hgfug , aa11lh , tnxtee |
| 2 | 1 | Over-specific event notification wording diverges from concise oracle-style message content | 8hgfug |
| 3 | 3 | Affirmative event-start email reply content is not canonicalized to the required acknowledgement structure | tnxtee |
| 4 | 4 | Multi-ambiguity user clarifications omit explicit questions for all unresolved decision dimensions | 2qw4nk , qrxrry |
| 5 | 4 | Time-sensitive repeated actions drift against deadlines and stop before the required count is complete | 71j6lf |
| 6 | 4 | Lookup tools reject raw identifiers that downstream action tools accept, causing resolution loops in deadline-bound tasks | 71j6lf |
| ID | Iter. | Induced Failure Pattern | Supporting Scenarios |
| 1 | 1 | Premature completion on constrained file edits without rule-aware final validation | overfull-hbox |
| 2 | 1 | Long-running probabilistic computation lacks staged dependency setup, diagnostic runs, and progress-aware output validation | mcmc-sampling-stan |
| 3 | 2 | Security sanitizer tasks finalize after weak smoke tests without adversarial and clean-preservation validation | filter-js-from-html |
| 4 | 4 | Binary extraction solutions finalize after superficial output smoke tests without independent address-space convention validation | extract-elf |
| 5 | 4 | Agent self-validation leaves stale runtime artifacts that change verifier timing and hide required fresh-run output | make-doom-for-mips |
| 6 | 4 | Completion-only validation gates do not help long build/debug tasks that time out before attempting finalization | make-doom-for-mips |