Organizations: Generative AI Lab, College of Computing and Data Science, Nanyang Technological University, Singapore 639798 · Tongyi Lab, Alibaba Group
Agentic Harness is the runtime that constructs task context and controls execution flow, thereby shaping overall agent performance. Given a fixed model and external evaluation, automated Harness optimization seeks to improve this runtime through an iterative proposal--evaluation loop to better solve target tasks. Existing meta-harness methods mainly adopt proposer-centric discovery, in which an LLM-based proposer integrates accumulated experimental findings to determine subsequent Harness revisions. This places the burden of maintaining the evolving improvement state on the proposer as history expands and its underlying experimental logic becomes harder to discern. In this paper, we introduce the Causal Improvement Graph (CIG), a graph-governed meta-harness framework that externalizes the evolving improvement state in a persistent graph, allowing prior findings to directly govern subsequent Harness optimization through local proposer operations. CIG grows and links Evidence, Hypothesis, Intervention, and Outcome nodes to represent what was observed, how it may be explained, how to test that explanation, and what the evaluation reveals. Their structural relations preserve how the improvement state changes across iterations, allowing local proposers to build directly on relations among prior findings rather than recover them from raw history. Across various agent tasks, CIG discovers stronger Harnesses than previous meta-harness baselines and remains robust to the choice of task solver and proposer. Structural ablations further support the design of an explicit improvement state with graph-governed evolution.
Figures & tables
Figure 1: Comparison of proposer-centric discovery and graph-governed discovery. Left: a proposer follows a prescribed procedure to manage an end-to-end process over interaction records and candidate artifacts. Right: CIG makes the growing graph the persistent improvement state, invokes proposers for locally scoped operations, and updates the graph and its assessments using externally evaluated outcomes.
Figure 2: Improvement state transitions in CIG. Left: graph-governed discovery operations transform Gt into Gt+1 by adding and linking evidence, hypothesis, intervention, and outcome nodes and revising existing hypothesis assessments based on evaluated outcomes. Right: typed records represent the improvement state and reference the corresponding Harness artifacts.
Method
SWE-V
AppWorld
TB2.0-V
Macro avg.
Published HarnessFix results (3-run means)
Initial Harness ( Chen et al., 2026b )
45.3
36.7
17.6
33.2
GEPA ( Agrawal et al., 2026 )
46.7
37.4
19.6
34.6
SCOPE ( Pei et al., 2025 )
48.3
38.9
20.6
35.9
ReCreate ( Hao et al., 2026 )
51.7
39.3
21.6
37.5
Meta-Harness ( Lee et al., 2026b )
54.7
40.4
23.5
39.5
Table 1: Held-out test performance of optimized Harnesses under the HarnessFix evaluation protocol. Task and optimization agents use the GPT-5 mini configuration ( OpenAI, 2025 ) . Entries are test task-completion rates (%), averaged over three runs. Published rows reproduce HarnessFix’s reported results, and our runs use the aligned evaluation setup. Full History is our baseline, matched to CIG in initialization and evaluation allowance. Parentheses in the CIG row show percentage-point gains over the reproduced initial Harness. Reporting conventions are given in Appendix A.8 .
Val-selected Test ↑
Δ Test ↑
Selection regret ↓
Solver
Full History
Wide
Deep
Wide
Deep
Full History
Wide
Deep
Qwen3.5-Plus
51.5
60.6
59.7
+9.1
+8.2
7.1
0.0
0.0
Qwen3.7-Flash
41.3
53.8
49.4
+12.5
+8.2
0.0
0.0
2.0
DeepSeek-v4-Flash
51.3
55.2
51.7
+3.9
+0.4
0.0
0.0
3.9
GLM-5.2-Fast-Preview
50.4
54.7
50.3
+4.3
-0.1
0.0
0.0
2.5
MiniMax-M2.5
45.4
48.9
47.8
+3.5
+2.4
0.0
0.2
0.1
Table 2: Robustness across task solvers on the three classification tasks. Entries are mean test scores (%) and selection regret under eight candidate evaluations. Δ reports test gains over Full History in percentage points. Wide uses two waves of four candidates, and Deep uses four waves of two.
Figure 3: Search progress and graph growth on LawBench using post-search validation re-evaluation means. (a) Best-so-far scores at completed wave boundaries. (b) Filled markers show individual candidate scores, linked to Outcome nodes; hollow markers and step curves show the best-so-far score, as in (a). Insets show node counts before and after updates, with gray marking retained nodes and colors marking additions. Iteration zero denotes the initial graph state before the first round of Harness optimization.
Proposer
Full History
CIG (Wide)
Δ
GPT-5.5
58.6 ±0.5
61.0 ±0.0
+2.4
GPT-5.6 Sol
59.0 ±1.0
64.0 ±1.0
+5.0
GLM-5.3
58.2 ±1.1
65.6 ±0.5
+7.4
Kimi K3
56.6 ±0.5
65.4 ±0.5
+8.8
DeepSeek-V4-Pro
59.0 ±0.7
63.0 ±1.0
+4.0
Table 3: Robustness across proposers on SWE-bench Verified. Scores are mean test pass@1 ± sample standard deviation of the validation-selected Harness. Δ is the difference in means.
Figure 4: Mean task-level test performance. Labels show CIG (Wide) scores and gains over Full History in percentage points. USPTO denotes USPTO-50k, DS denotes DeepSeek, and Pre denotes Preview. Each panel uses its own linear axis with a nonzero origin. Bar lengths are comparable within each panel, but not across panels with different axis ranges.
Cumulative construction
Proposer A
Proposer B
Variant
State
Relations
Update
Selection
Test ↑
AUC ↑
Test ↑
AUC ↑
M0 (Full History)
Raw
None
Reset
Indep.
58.0
58.0
56.4
56.8
M1 (+ Typed state)
Typed
None
Reset
Indep.
59.6
(+1.6)
59.1
57.4
(+1.0)
57.9
M2 (+ Relations)
Typed
Correct
Reset
Indep.
60.8
(+1.2)
60.5
58.8
(+1.4)
59.2
↪ B1 (Shuffled)
Typed
Shuffled
Reset
Indep.
59.0
(-1.8)
59.2
57.4
(-1.4)
58.1
M3 (+ Revision)
Typed
Correct
Revise
Indep.
62.2
(+1.4)
61.8
59.6
(+0.8)
60.4
Table 4: Structural contributions to CIG. Starting from Full History, cumulative variants introduce typed state, relations, and revision before full CIG, while alternative designs (blue) test how these components are used. The results support their complementary value for final Harness quality and search progress under a matched workflow. Appendix A.3 provides exact variant definitions and reporting conventions.
Appendix figures & tables4 assets
Supplementary material from the paper’s appendix.
Appendix
Method
Persistent memory
Explicit predictions
Graph state
Compared with CIG
ADAS ( Hu et al., 2025 )
✓
–
–
Organizes search around archived programs, while CIG organizes it around revisable hypotheses.
AFlow ( Zhang et al., 2025b )
✓
–
✓
Searches a workflow tree, while CIG links evidence and hypotheses across interventions.
Meta-Harness ( Lee et al., 2026b )
✓
–
–
A proposer interprets stored history, while CIG exposes linked assessments to scoped operations.
AHE ( Lin et al., 2026 )
✓
✓
–
Checks edit-level predictions, while CIG carries their assessments into a persistent hypothesis graph.
HarnessFix ( Chen et al., 2026b )
✓
✓
–
Organizes repair memory, while CIG uses linked hypothesis assessments to select experiments.
Living-Harness ( Du et al., 2026 )
✓
–
✓
Updates procedural behavior, while CIG updates the optimizer’s experimental reasoning state.
Appendix
Table 5: State representations in related methods. Columns indicate persistent search or experience memory, explicit predictions or hypotheses, and graph-structured state used in search or adaptation. The final column contrasts each method’s organizing principle with CIG. ✓ denotes an explicitly described feature, and – denotes a feature not established by the cited description.
Variant
GPT-5.6 Sol
DeepSeek-V4-Pro
M0 (Full History)
58.0 ±0.0
56.4 ±0.9
M1 (+ Typed State)
59.6 ±0.5
57.4 ±0.5
M2 (+ Logic Relations)
60.8 ±0.8
58.8 ±0.4
B1 (Shuffled Relations)
59.0 ±0.0
57.4 ±0.5
M3 (+ Revision-aware Persistence)
62.2 ±0.8
59.6 ±0.9
B2 (Append-only Persistence)
61.0 ±0.0
58.8 ±0.4
Appendix
Table 6: Companion statistics for Table 4 . Test scores are mean ± sample standard deviation in percentage points, grouped by proposer.
Method / checkpoint
Tokens (M)
Ratio
Candidates
Test (%)
Full History (complete)
0.530
1.00×
8
41.3
CIG Wide (budget checkpoint)
0.530
1.00×
5
46.8
CIG Wide (complete)
0.820
1.55×
8
52.9
Appendix
Table 7: Proposer token usage and validation-selected test performance in a representative configuration. Token ratios use the complete Full History run as reference.