Organizations: Generative AI Lab, College of Computing and Data Science, Nanyang Technological University, Singapore 639798 · Tongyi Lab, Alibaba Group
Agentic Harness is the runtime that constructs task context and controls execution flow, thereby shaping overall agent performance. Given a fixed model and external evaluation, automated Harness optimization seeks to improve this runtime through an iterative proposal--evaluation loop to better solve target tasks. Existing meta-harness methods mainly adopt proposer-centric discovery, in which an LLM-based proposer integrates accumulated experimental findings to determine subsequent Harness revisions. This places the burden of maintaining the evolving improvement state on the proposer as history expands and its underlying experimental logic becomes harder to discern. In this paper, we introduce the Causal Improvement Graph (CIG), a graph-governed meta-harness framework that externalizes the evolving improvement state in a persistent graph, allowing prior findings to directly govern subsequent Harness optimization through local proposer operations. CIG grows and links Evidence, Hypothesis, Intervention, and Outcome nodes to represent what was observed, how it may be explained, how to test that explanation, and what the evaluation reveals. Their structural relations preserve how the improvement state changes across iterations, allowing local proposers to build directly on relations among prior findings rather than recover them from raw history. Across various agent tasks, CIG discovers stronger Harnesses than previous meta-harness baselines and remains robust to the choice of task solver and proposer. Structural ablations further support the design of an explicit improvement state with graph-governed evolution.
Figures & tables
Figure 1: Comparison of proposer-centric discovery and graph-governed discovery. Left: a proposer follows a prescribed procedure to manage an end-to-end process over interaction records and candidate artifacts. Right: CIG makes the growing graph the persistent improvement state, invokes proposers for locally scoped operations, and updates the graph and its assessments using externally evaluated outcomes.
Figure 2: Improvement state transitions in CIG. Left: graph-governed discovery operations transform Gt into Gt+1 by adding and linking evidence, hypothesis, intervention, and outcome nodes and revising existing hypothesis assessments based on evaluated outcomes. Right: typed records represent the improvement state and reference the corresponding Harness artifacts.
Method
SWE-V
AppWorld
TB2.0-V
Macro avg.
Published HarnessFix results (3-run means)
Initial Harness ( Chen et al., 2026b )
45.3
36.7
17.6
33.2
GEPA ( Agrawal et al., 2026 )
46.7
37.4
19.6
34.6
SCOPE ( Pei et al., 2025 )
48.3
38.9
20.6
35.9
ReCreate ( Hao et al., 2026 )
51.7
39.3
21.6
37.5
Meta-Harness ( Lee et al., 2026b )
54.7
40.4
23.5
39.5
Table 1: Held-out test performance of optimized Harnesses under the HarnessFix evaluation protocol. Task and optimization agents use the GPT-5 mini configuration ( OpenAI, 2025 ) . Entries are test task-completion rates (%), averaged over three runs. Published rows reproduce HarnessFix’s reported results, and our runs use the aligned evaluation setup. Full History is our baseline, matched to CIG in initialization and evaluation allowance. Parentheses in the CIG row show percentage-point gains over the reproduced initial Harness. Reporting conventions are given in Appendix A.8 .
Val-selected Test ↑
Δ Test ↑
Selection regret ↓
Solver
Full History
Wide
Deep
Wide
Deep
Full History
Wide
Deep
Qwen3.5-Plus
51.5
60.6
59.7
+9.1
+8.2
7.1
0.0
0.0
Qwen3.7-Flash
41.3
53.8
49.4
+12.5
+8.2
0.0
0.0
2.0
DeepSeek-v4-Flash
51.3
55.2
51.7
+3.9
+0.4
0.0
0.0
3.9
GLM-5.2-Fast-Preview
50.4
54.7
50.3
+4.3
-0.1
0.0
0.0
2.5
MiniMax-M2.5
45.4
48.9
47.8
+3.5
+2.4
0.0
0.2
0.1
Table 2: Robustness across task solvers on the three classification tasks. Entries are mean test scores (%) and selection regret under eight candidate evaluations. Δ reports test gains over Full History in percentage points. Wide uses two waves of four candidates, and Deep uses four waves of two.
Figure 3: Search progress and graph growth on LawBench using post-search validation re-evaluation means. (a) Best-so-far scores at completed wave boundaries. (b) Filled markers show individual candidate scores, linked to Outcome nodes; hollow markers and step curves show the best-so-far score, as in (a). Insets show node counts before and after updates, with gray marking retained nodes and colors marking additions. Iteration zero denotes the initial graph state before the first round of Harness optimization.
Proposer
Full History
CIG (Wide)
Δ
GPT-5.5
58.6 ±0.5
61.0 ±0.0
+2.4
GPT-5.6 Sol
59.0 ±1.0
64.0 ±1.0
+5.0
GLM-5.3
58.2 ±1.1
65.6 ±0.5
+7.4
Kimi K3
56.6 ±0.5
65.4 ±0.5
+8.8
DeepSeek-V4-Pro
59.0 ±0.7
63.0 ±1.0
+4.0
Table 3: Robustness across proposers on SWE-bench Verified. Scores are mean test pass@1 ± sample standard deviation of the validation-selected Harness. Δ is the difference in means.
Figure 4: Mean task-level test performance. Labels show CIG (Wide) scores and gains over Full History in percentage points. USPTO denotes USPTO-50k, DS denotes DeepSeek, and Pre denotes Preview. Each panel uses its own linear axis with a nonzero origin. Bar lengths are comparable within each panel, but not across panels with different axis ranges.
Cumulative construction
Proposer A
Proposer B
Variant
State
Relations
Update
Selection
Test ↑
AUC ↑
Test ↑
AUC ↑
M0 (Full History)
Raw
None
Reset
Indep.
58.0
58.0
56.4
56.8
M1 (+ Typed state)
Typed
None
Reset
Indep.
59.6
(+1.6)
59.1
57.4
(+1.0)
57.9
M2 (+ Relations)
Typed
Correct
Reset
Indep.
60.8
(+1.2)
60.5
58.8
(+1.4)
59.2
↪ B1 (Shuffled)
Typed
Shuffled
Reset
Indep.
59.0
(-1.8)
59.2
57.4
(-1.4)
58.1
M3 (+ Revision)
Typed
Correct
Revise
Indep.
62.2
(+1.4)
61.8
59.6
(+0.8)
60.4
Table 4: Structural contributions to CIG. Starting from Full History, cumulative variants introduce typed state, relations, and revision before full CIG, while alternative designs (blue) test how these components are used. The results support their complementary value for final Harness quality and search progress under a matched workflow. Appendix A.3 provides exact variant definitions and reporting conventions.
Appendix figures & tables4 assets
Supplementary material from the paper’s appendix.
Appendix
Method
Persistent memory
Explicit predictions
Graph state
Compared with CIG
ADAS ( Hu et al., 2025 )
✓
–
–
Organizes search around archived programs, while CIG organizes it around revisable hypotheses.
AFlow ( Zhang et al., 2025b )
✓
–
✓
Searches a workflow tree, while CIG links evidence and hypotheses across interventions.
Meta-Harness ( Lee et al., 2026b )
✓
–
–
A proposer interprets stored history, while CIG exposes linked assessments to scoped operations.
AHE ( Lin et al., 2026 )
✓
✓
–
Checks edit-level predictions, while CIG carries their assessments into a persistent hypothesis graph.
HarnessFix ( Chen et al., 2026b )
✓
✓
–
Organizes repair memory, while CIG uses linked hypothesis assessments to select experiments.
Living-Harness ( Du et al., 2026 )
✓
–
✓
Updates procedural behavior, while CIG updates the optimizer’s experimental reasoning state.
Appendix
Table 5: State representations in related methods. Columns indicate persistent search or experience memory, explicit predictions or hypotheses, and graph-structured state used in search or adaptation. The final column contrasts each method’s organizing principle with CIG. ✓ denotes an explicitly described feature, and – denotes a feature not established by the cited description.
Variant
GPT-5.6 Sol
DeepSeek-V4-Pro
M0 (Full History)
58.0 ±0.0
56.4 ±0.9
M1 (+ Typed State)
59.6 ±0.5
57.4 ±0.5
M2 (+ Logic Relations)
60.8 ±0.8
58.8 ±0.4
B1 (Shuffled Relations)
59.0 ±0.0
57.4 ±0.5
M3 (+ Revision-aware Persistence)
62.2 ±0.8
59.6 ±0.9
B2 (Append-only Persistence)
61.0 ±0.0
58.8 ±0.4
Appendix
Table 6: Companion statistics for Table 4 . Test scores are mean ± sample standard deviation in percentage points, grouped by proposer.
Method / checkpoint
Tokens (M)
Ratio
Candidates
Test (%)
Full History (complete)
0.530
1.00×
8
41.3
CIG Wide (budget checkpoint)
0.530
1.00×
5
46.8
CIG Wide (complete)
0.820
1.55×
8
52.9
Appendix
Table 7: Proposer token usage and validation-selected test performance in a representative configuration. Token ratios use the complete Full History run as reference.
AI agent performance depends critically on the runtime harness, comprising the prompts, tools, memory, and control flow that mediate how a model observes, reasons, and acts. Yet today's harnesses remain largely hand-crafted and static: each new model or task still demands bespoke scaffolding, and the rich traces produced during execution are rarely distilled back into systematic improvement. We introduce HarnessX, a foundry for composable, adaptive, and evolvable agent harnesses. HarnessX assembles typed harness primitives via a substitution algebra, adapts them through AEGIS, a trace-driven multi-agent evolution engine grounded in an operational mirror between symbolic adaptation and reinforcement learning, and closes the harness-model loop by turning trajectories into both harness updates and model training signal. Across five benchmarks (ALFWorld, GAIA, WebShop, tau^3-Bench, and SWE-bench Verified), HarnessX yields an average gain of +14.5% (up to +44.0%), with gains largest where baselines are lowest. These results suggest that agent progress need not come from model scaling alone: composing and evolving runtime interfaces from execution feedback is an actionable and complementary lever. The complete codebase will be open-sourced in a future release.
An agent harness is the code that organizes context, maintains state, and coordinates tool calls for a language model. We study how to improve the harness under a limited evaluation budget while keeping model weights fixed. Our method, MESH-Harness, organizes each harness into functional modules with explicit role-specific interfaces, allowing alternative implementations of each module to be substituted and recombined. It uses shared module representations and full-covariance LinUCB to score candidate combinations based on predicted performance and exploration value. Mixed-start coordinate ascent selects complete configurations for evaluation without enumerating the combinatorial space. Validation traces then guide local code edits, and the resulting candidates are incorporated into fixed-capacity role-specific pools for subsequent recombination. On text tasks, retrieval-augmented mathematical reasoning, code generation, and interactive scientific tasks, MESH-Harness outperforms Meta-Harness by 5.70, 7.01, 2.00, and 5.00 points, respectively, under matched candidate-evaluation budgets. Iterative harness optimization improves MESH-Harness by 5.63-7.79 points over its first-round configurations. For the reported configurations, aggregate test-time cost is 44.2% lower than that of Meta-Harness, while total cost including search is 14.6% lower. These results show that combining module-level design reuse with feedback-driven compositional search can systematically improve agent harnesses while keeping overall optimization cost under control.
Zhiwei Shang, Yu Huo, Mingrong Gong +6
The Chinese University of Hong Kong, Shenzhen · DeepWisdom · The Chinese University of Hong Kong +3
Under model--harness co-evolution, harnesses are not merely inference-time scaffolds but data-generating components whose execution traces can shape future foundation models. This motivates harness-in-the-loop learning: optimizing harnesses for both immediate agent performance and the quality of traces used for future model training. However, continually updating provider-built scaffolds is costly and labor-intensive. We therefore investigate whether optimizing user-constructed harnesses in a task-specific manner can improve execution-trace quality while remaining computationally lightweight and requiring only a few update iterations. To this end, we introduce Recursive Harness Self-Improvement (RHI), which represents the harness as a prompt-level specification of the agent loop and iteratively refines it using pairwise feedback over its own revision history. Across 30 synthetic machine-learning research tasks spanning quantitative finance, robotics, and pharmacy, a few RHI iterations suffice to substantially raise the performance ceiling of low-reasoning-effort agents, exceeding the corresponding maximum-reasoning-effort setting while reducing inference cost by up to 60%. We show that these gains arise primarily from improved task-specific context management through more effective inter-agent information flow rather than longer reasoning traces. Finally, we formalize this behavior as an information-theoretic hypothesis for RHI's implicit optimization objective, suggesting RHI as a practical algorithm for continual learning within the paradigm of model--harness co-evolution.