cs.AIOct 1, 2026

ActiveSaddler: Automated Curriculum Learning for Agent Harness Optimization

Authors: Sungho Park, Wonjoong Kim, Jue Zhang, Wook-Shin Han, Pengfei Gao, Chanyoung Park, Yongqiang Yao, Rao Fu, +3 more

Organizations: POSTECH · KAIST · Microsoft

Abstract

Automated harness optimization can substantially improve LLM agents by iteratively updating their prompts, tool interfaces, and control logic from execution feedback. However, existing methods primarily optimize how the harness is updated while largely fixing which training scenarios generate the feedback that drives those updates. As the harness evolves, the scenarios most useful for further optimization can change, suggesting that the training curriculum itself should adapt alongside the harness. We formulate this missing dimension of harness optimization as an automated curriculum learning problem and introduce ActiveSaddler. ActiveSaddler models the evolving curriculum as a non-stationary bandit with dynamically instantiated optimization targets. It abstracts recurring failures into reusable failure-pattern arms, estimates the potential learning progress from further targeting each pattern, and adaptively balances revisiting known weaknesses with exploring unseen scenarios for new ones. Optimization outcomes continually update both the set of discovered failure patterns and their priorities, allowing the curriculum to co-evolve with the harness. Experiments on GAIA2 and Terminal-Bench 2.0 show that ActiveSaddler consistently discovers stronger harnesses, improving test Pass@1 by 4.4 and 7.5 percentage points over the same harness optimizer using a scenario order fixed before optimization, respectively. Ablations further show that these gains depend on dynamically constructing optimization targets, estimating their evolving utility, and balancing continued optimization with new failure discovery. Together, these results establish automated curriculum learning as a new crucial optimization dimension for harness optimization.

Figures & tables

Appendix figures & tables18 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. HarnessCompass: Guiding Automatic Harness Evolution toward Generalizable and Effective Agent Harnesses

    Aug 3, 2026Luan Zhang, Ruochen Zhou, Dandan Song +9Evolving HarnessAgent Harness

  2. HarnessOpt-Bench: Evaluating LLMs at Harness Optimization

    Aug 6, 2026Varun Ursekar, Apaar Shanker, Yash Maurya +4Agent HarnessAgentic Benchmarks

  3. Co-Harness: Co-Evolving Harnesses and Model Weights for LLM Agents

    Jul 17, 2026Zhengyu Chen, Teng Xiao, Huaisheng Zhu +3Agent HarnessLarge Language Model Agents