Organizations: The Chinese University of Hong Kong, Shenzhen · DeepWisdom · The Chinese University of Hong Kong · Nanyang Technological University · Fudan University · National University of Singapore
An agent harness is the code that organizes context, maintains state, and coordinates tool calls for a language model. We study how to improve the harness under a limited evaluation budget while keeping model weights fixed. Our method, MESH-Harness, organizes each harness into functional modules with explicit role-specific interfaces, allowing alternative implementations of each module to be substituted and recombined. It uses shared module representations and full-covariance LinUCB to score candidate combinations based on predicted performance and exploration value. Mixed-start coordinate ascent selects complete configurations for evaluation without enumerating the combinatorial space. Validation traces then guide local code edits, and the resulting candidates are incorporated into fixed-capacity role-specific pools for subsequent recombination. On text tasks, retrieval-augmented mathematical reasoning, code generation, and interactive scientific tasks, MESH-Harness outperforms Meta-Harness by 5.70, 7.01, 2.00, and 5.00 points, respectively, under matched candidate-evaluation budgets. Iterative harness optimization improves MESH-Harness by 5.63-7.79 points over its first-round configurations. For the reported configurations, aggregate test-time cost is 44.2% lower than that of Meta-Harness, while total cost including search is 14.6% lower. These results show that combining module-level design reuse with feedback-driven compositional search can systematically improve agent harnesses while keeping overall optimization cost under control.
Figures & tables
Figure 1: Overview of MESH-Harness. (a) Each functional slot has a fixed interface and a pool of LLM-initialized implementations with cached embeddings . (b) Full-covariance LinUCB scores configurations with one implementation per slot , and mixed-start coordinate ascent selects a configuration for evaluation. (c) The selected harness is executed to obtain validation rewards and traces. (d) Trace diagnosis guides edits to implementations and fixed-capacity pool updates. Rewards update the bandit model ; new implementations join their respective pools for subsequent composition .
Figure 2: Test performance across four task domains. (a) Text macro-average, overall Math pass@1, LiveCodeBench solve rate under the fixed cap-8 protocol, and ScienceWorld mean progress score. (b) Unweighted mean of the four task scores. All scores are on a 0–100 scale, with higher values indicating better performance. MESH-Harness achieves the highest score on all four aggregate metrics.
Method
Text
Math
LiveCodeBench
ScienceWorld
SkillOpt
53.27
46.40
57.00
53.00
DGM-H
42.78
53.00
72.00
52.20
ADIAS
42.50
52.60
69.00
43.10
AHE
47.39
50.60
69.00
49.00
Meta-Harness
50.63
46.20
75.00
53.10
MESH-Harness
56.33
53.21
77.00
58.10
Table 1: Held-out test performance. All methods use 40 candidate evaluations; scores average three test runs. Text: dataset macro-average; Math: pooled pass@1; LiveCodeBench: cap-8 solve rate; ScienceWorld: mean progress. Scores use a 0–100 scale. Bold marks the column maximum; Δ compares MESH-Harness with Meta-Harness.
Harness
Score
No memory
26.07
Few-shot (8)
41.00
Few-shot (32)
46.87
Few-shot (all)
50.00
ACE memory
50.00
MCE memory
48.93
Table 2: Text reference harnesses. Fixed Meta-Harness seeds appear above the divider; scores are macro-averages.
Harness
Score
No retrieval
23.20
Random few-shot
41.42
BM25
42.18
ACE
42.38
Meta-Harness
46.20
MESH-Harness
53.21
Table 3: Math reference harnesses. Fixed Meta-Harness seeds appear above the divider; scores are overall pass@1.
Figure 3: Final performance relative to method-specific references. “Initial” denotes a fixed reference seed for Meta-Harness and the first evaluated composition for MESH-Harness, not a shared initialization. Meta-Harness references are 50.00 (Text), 42.38 (Math, ACE), 71.00 (LiveCodeBench), and 55.50 (ScienceWorld). Brackets show final minus reference.
Figure 4: Exploration strength. Text macro-average versus α ; the star marks the default. Reference lines show Meta-Harness and random selection.
Selection
Score
Δ
No UCB bonus ( α=0 )
51.70
−4.63
α=0.2
53.17
−3.16
α=0.4
54.87
−1.46
Full ( α=0.8 )
56.33
0.00
α=1.6
55.10
−1.23
Random composition
48.80
−7.53
Table 4: Configuration selection. Text macro-average; Δ is relative to full MESH-Harness. Rows follow the α sweep at left.
Variant
Score
Δ
MESH-Harness
56.33
0.00
No pool refresh
51.20
−5.13
Refresh without historical traces
53.80
−2.53
Meta-Harness
50.63
−5.70
Table 5: Module evolution. Text macro-average over three test runs and change from the full method.
Search + test
Test only
Task
Meta
MESH
Saving
Meta
MESH
Saving
Text
63.1843
51.1643
19.0%
0.1709
0.1662
2.8%
Math
58.8652
47.2431
19.7%
0.5475
0.4987
8.9%
LiveCodeBench
37.1387
34.7826
6.3%
0.5273
0.3847
27.0%
ScienceWorld
35.8655
33.4178
6.8%
2.3071
0.9339
59.5%
Total
195.0537
166.6078
14.6%
3.5528
1.9835
44.2%
Table 6: Search and execution costs (USD). Search + test includes optimization and testing the selected configurations; test only excludes optimization. Savings are relative to Meta-Harness. Totals sum the four reported task workloads.
Figure 5: Score–cost tradeoff. Circles: MESH-Harness; squares: Meta-Harness. The dashed line connects nondominated evaluated configurations. Cost uses a log scale.
Score ↑
Cost ↓
Execution model
Meta
MESH
Meta
MESH
DeepSeek V4 Flash
50.63
56.33
0.1709
0.1662
GPT-5.6 Luna
51.00
55.37
0.5359
0.4987
Qwen 3.8 Flash
49.17
53.70
0.2608
0.2496
Claude Sonnet 5
55.57
59.93
7.4215
6.8874
Table 7: Frozen-harness transfer. Text macro-average and execution cost (USD), without further search. Bold marks the better result within each model.
Dataset
Meta-Harness solution
+ MESH refinement
LawBench
51.00
55.00
Symptom-to-Diagnosis
84.90
90.10
USPTO
16.00
18.00
Macro-average
50.63
54.37
Table 8: Refining a Meta-Harness solution. Held-out text scores before and after 20 additional MESH-Harness candidate evaluations.
Appendix figures & tables3 assets
Supplementary material from the paper’s appendix.
Appendix
Dataset
Train
Validation
Test
Prediction or interaction unit
USPTO
50
30
100
Precursor strings for a product molecule
Symptom-to-Diagnosis
200
50
212
A diagnosis from a symptom description
LawBench
200
50
100
A set of charges from a case description
Mathematical reasoning
–
400
500
A final mathematical answer
LiveCodeBench
100
100
100
One selected Python program per problem
ScienceWorld
66
22
102
One interactive episode
Appendix
Table 9: Task splits and prediction units. A dash indicates no separate training split. Math contains 200 validation problems from each source; its 500 test problems comprise 254 from OlympiadBench and 246 from MathArena. ScienceWorld counts are episodes.
Retrieval source
Problem–solution pairs
NuminaMath
55,968
Omni-MATH
3,201
DeepMath
98,670
OpenMathReasoning
267,647
PolyMath
8,440
Total
433,926
Appendix
Table 10: Math retrieval corpus after evaluation-overlap filtering and deduplication.
Setting
Value
Candidate-evaluation budget
40
Implementations per slot
10
Feature dimensions
d=128 per slot; p=896 for seven slots, p=640 for five
Ridge regularization; UCB coefficient
λ=1 ; α=0.8
Coverage; family-reuse penalty
η=0.04 ; γ=0.04
Similarity attenuation; minimum scale
β=0.25 ; gmin=0.5
Appendix
Table 11: Default search settings. The window W and refresh interval are measured in completed candidate evaluations.