Organizations: The Chinese University of Hong Kong, Shenzhen · DeepWisdom · The Chinese University of Hong Kong · Nanyang Technological University · Fudan University · National University of Singapore
An agent harness is the code that organizes context, maintains state, and coordinates tool calls for a language model. We study how to improve the harness under a limited evaluation budget while keeping model weights fixed. Our method, MESH-Harness, organizes each harness into functional modules with explicit role-specific interfaces, allowing alternative implementations of each module to be substituted and recombined. It uses shared module representations and full-covariance LinUCB to score candidate combinations based on predicted performance and exploration value. Mixed-start coordinate ascent selects complete configurations for evaluation without enumerating the combinatorial space. Validation traces then guide local code edits, and the resulting candidates are incorporated into fixed-capacity role-specific pools for subsequent recombination. On text tasks, retrieval-augmented mathematical reasoning, code generation, and interactive scientific tasks, MESH-Harness outperforms Meta-Harness by 5.70, 7.01, 2.00, and 5.00 points, respectively, under matched candidate-evaluation budgets. Iterative harness optimization improves MESH-Harness by 5.63-7.79 points over its first-round configurations. For the reported configurations, aggregate test-time cost is 44.2% lower than that of Meta-Harness, while total cost including search is 14.6% lower. These results show that combining module-level design reuse with feedback-driven compositional search can systematically improve agent harnesses while keeping overall optimization cost under control.
Figures & tables
Figure 1: Overview of MESH-Harness. (a) Each functional slot has a fixed interface and a pool of LLM-initialized implementations with cached embeddings . (b) Full-covariance LinUCB scores configurations with one implementation per slot , and mixed-start coordinate ascent selects a configuration for evaluation. (c) The selected harness is executed to obtain validation rewards and traces. (d) Trace diagnosis guides edits to implementations and fixed-capacity pool updates. Rewards update the bandit model ; new implementations join their respective pools for subsequent composition .
Figure 2: Test performance across four task domains. (a) Text macro-average, overall Math pass@1, LiveCodeBench solve rate under the fixed cap-8 protocol, and ScienceWorld mean progress score. (b) Unweighted mean of the four task scores. All scores are on a 0–100 scale, with higher values indicating better performance. MESH-Harness achieves the highest score on all four aggregate metrics.
Method
Text
Math
LiveCodeBench
ScienceWorld
SkillOpt
53.27
46.40
57.00
53.00
DGM-H
42.78
53.00
72.00
52.20
ADIAS
42.50
52.60
69.00
43.10
AHE
47.39
50.60
69.00
49.00
Meta-Harness
50.63
46.20
75.00
53.10
MESH-Harness
56.33
53.21
77.00
58.10
Table 1: Held-out test performance. All methods use 40 candidate evaluations; scores average three test runs. Text: dataset macro-average; Math: pooled pass@1; LiveCodeBench: cap-8 solve rate; ScienceWorld: mean progress. Scores use a 0–100 scale. Bold marks the column maximum; Δ compares MESH-Harness with Meta-Harness.
Harness
Score
No memory
26.07
Few-shot (8)
41.00
Few-shot (32)
46.87
Few-shot (all)
50.00
ACE memory
50.00
MCE memory
48.93
Table 2: Text reference harnesses. Fixed Meta-Harness seeds appear above the divider; scores are macro-averages.
Harness
Score
No retrieval
23.20
Random few-shot
41.42
BM25
42.18
ACE
42.38
Meta-Harness
46.20
MESH-Harness
53.21
Table 3: Math reference harnesses. Fixed Meta-Harness seeds appear above the divider; scores are overall pass@1.
Figure 3: Final performance relative to method-specific references. “Initial” denotes a fixed reference seed for Meta-Harness and the first evaluated composition for MESH-Harness, not a shared initialization. Meta-Harness references are 50.00 (Text), 42.38 (Math, ACE), 71.00 (LiveCodeBench), and 55.50 (ScienceWorld). Brackets show final minus reference.
Figure 4: Exploration strength. Text macro-average versus α ; the star marks the default. Reference lines show Meta-Harness and random selection.
Selection
Score
Δ
No UCB bonus ( α=0 )
51.70
−4.63
α=0.2
53.17
−3.16
α=0.4
54.87
−1.46
Full ( α=0.8 )
56.33
0.00
α=1.6
55.10
−1.23
Random composition
48.80
−7.53
Table 4: Configuration selection. Text macro-average; Δ is relative to full MESH-Harness. Rows follow the α sweep at left.
Variant
Score
Δ
MESH-Harness
56.33
0.00
No pool refresh
51.20
−5.13
Refresh without historical traces
53.80
−2.53
Meta-Harness
50.63
−5.70
Table 5: Module evolution. Text macro-average over three test runs and change from the full method.
Search + test
Test only
Task
Meta
MESH
Saving
Meta
MESH
Saving
Text
63.1843
51.1643
19.0%
0.1709
0.1662
2.8%
Math
58.8652
47.2431
19.7%
0.5475
0.4987
8.9%
LiveCodeBench
37.1387
34.7826
6.3%
0.5273
0.3847
27.0%
ScienceWorld
35.8655
33.4178
6.8%
2.3071
0.9339
59.5%
Total
195.0537
166.6078
14.6%
3.5528
1.9835
44.2%
Table 6: Search and execution costs (USD). Search + test includes optimization and testing the selected configurations; test only excludes optimization. Savings are relative to Meta-Harness. Totals sum the four reported task workloads.
Figure 5: Score–cost tradeoff. Circles: MESH-Harness; squares: Meta-Harness. The dashed line connects nondominated evaluated configurations. Cost uses a log scale.
Score ↑
Cost ↓
Execution model
Meta
MESH
Meta
MESH
DeepSeek V4 Flash
50.63
56.33
0.1709
0.1662
GPT-5.6 Luna
51.00
55.37
0.5359
0.4987
Qwen 3.8 Flash
49.17
53.70
0.2608
0.2496
Claude Sonnet 5
55.57
59.93
7.4215
6.8874
Table 7: Frozen-harness transfer. Text macro-average and execution cost (USD), without further search. Bold marks the better result within each model.
Dataset
Meta-Harness solution
+ MESH refinement
LawBench
51.00
55.00
Symptom-to-Diagnosis
84.90
90.10
USPTO
16.00
18.00
Macro-average
50.63
54.37
Table 8: Refining a Meta-Harness solution. Held-out text scores before and after 20 additional MESH-Harness candidate evaluations.
Appendix figures & tables3 assets
Supplementary material from the paper’s appendix.
Appendix
Dataset
Train
Validation
Test
Prediction or interaction unit
USPTO
50
30
100
Precursor strings for a product molecule
Symptom-to-Diagnosis
200
50
212
A diagnosis from a symptom description
LawBench
200
50
100
A set of charges from a case description
Mathematical reasoning
–
400
500
A final mathematical answer
LiveCodeBench
100
100
100
One selected Python program per problem
ScienceWorld
66
22
102
One interactive episode
Appendix
Table 9: Task splits and prediction units. A dash indicates no separate training split. Math contains 200 validation problems from each source; its 500 test problems comprise 254 from OlympiadBench and 246 from MathArena. ScienceWorld counts are episodes.
Retrieval source
Problem–solution pairs
NuminaMath
55,968
Omni-MATH
3,201
DeepMath
98,670
OpenMathReasoning
267,647
PolyMath
8,440
Total
433,926
Appendix
Table 10: Math retrieval corpus after evaluation-overlap filtering and deduplication.
Setting
Value
Candidate-evaluation budget
40
Implementations per slot
10
Feature dimensions
d=128 per slot; p=896 for seven slots, p=640 for five
Ridge regularization; UCB coefficient
λ=1 ; α=0.8
Coverage; family-reuse penalty
η=0.04 ; γ=0.04
Similarity attenuation; minimum scale
β=0.25 ; gmin=0.5
Appendix
Table 11: Default search settings. The window W and refresh interval are measured in completed candidate evaluations.
AI agent performance depends critically on the runtime harness, comprising the prompts, tools, memory, and control flow that mediate how a model observes, reasons, and acts. Yet today's harnesses remain largely hand-crafted and static: each new model or task still demands bespoke scaffolding, and the rich traces produced during execution are rarely distilled back into systematic improvement. We introduce HarnessX, a foundry for composable, adaptive, and evolvable agent harnesses. HarnessX assembles typed harness primitives via a substitution algebra, adapts them through AEGIS, a trace-driven multi-agent evolution engine grounded in an operational mirror between symbolic adaptation and reinforcement learning, and closes the harness-model loop by turning trajectories into both harness updates and model training signal. Across five benchmarks (ALFWorld, GAIA, WebShop, tau^3-Bench, and SWE-bench Verified), HarnessX yields an average gain of +14.5% (up to +44.0%), with gains largest where baselines are lowest. These results suggest that agent progress need not come from model scaling alone: composing and evolving runtime interfaces from execution feedback is an actionable and complementary lever. The complete codebase will be open-sourced in a future release.
Long-horizon language-model agents are dominated, in lines of code and in operational complexity, not by their underlying model but by the harness that wraps it: context compaction, tool caching, semantic memory, trajectory reuse, speculative tool prediction, and the glue that binds the model to a sandboxed execution environment. We argue that harness design is a first-class machine-learning problem and that automated configuration search dominates manual stacking once the flag space exceeds a handful of bits. We defend this claim in two steps. First, we formalize automated harness optimization as constrained noisy Bayesian optimization over a mixed-variable, cost-heterogeneous configuration space with cold-start-corrected rewards and a posterior chance-constrained safety check, and give a reference solver, HARBOR (Harness Axis-aligned Regularized Bayesian Optimization Routine), built from a block-additive SAAS surrogate, multi-fidelity cost-aware acquisition, and TuRBO trust regions. Second, we instantiate the problem in a flag-gated harness over a production coding agent and report a controlled four-round manual-tuning case study against a fixed task suite and an end-to-end HARBOR run. The formulation itself is task-class agnostic: the configuration space, reward correction, acquisition, and safety check apply to any agent harness with a bounded flag space and a reproducible task suite.
Harness design plays a critical role in agent performance by shaping how large language models (LLMs) perceive, reason over, and act within executable environments. Recent work has proposed automatic harness evolution, which iteratively improves the harness from agent--environment interactions. However, existing methods often overfit to the evolution tasks, rely exclusively on trajectory-derived signals, and optimize harness components jointly, causing interference across components. We propose HarnessCompass, a novel automatic harness evolution framework built around constrained evolution, proactive feedback, and component-wise optimization. HarnessCompass first enforces global constraints on evolution, restricting modifications to task-agnostic harness changes that generalize beyond the evolution tasks. It then augments trajectory-derived evidence with proactive first-person feedback from the agent about harness usage, yielding richer signals for evolution. Finally, it decouples the optimization of different harness components before consolidating them into a unified harness, reducing cross-component interference while preserving component synergy. On SWE-bench Verified with GPT-5.4, HarnessCompass improves Pass@1 from 54% to 66% in only 5 evolution iterations, outperforming AHE in both effectiveness and evolution efficiency. In addition, the evolved harness transfers effectively to held-out tasks and other models, demonstrating substantially stronger generalization than prior automatic harness evolution methods.
Luan Zhang, Ruochen Zhou, Dandan Song +9
Beijing Institute of Technology, China · City University of Hong Kong, China · Independent, China