Agent harnesses govern how large language models (LLMs) gather context, invoke tools, verify results, preserve state, and terminate, largely affecting agent performance. However, the value of each harness mechanism can differ across heterogeneous tasks: a mechanism that improves one task may impose overhead or context distraction on another, leading to the suboptimality of a global harness. We characterize this suboptimality as a mismatch induced by fixed mechanism choices, motivating task-specific harness construction. Nonetheless, generating harness code for each task introduces generation and debugging costs, with execution risks that can compound as more mechanisms are generated. To address those challenges, we introduce Harness Primitives, reusable harness mechanisms with clear application scope and composition contract mined from failed task trajectories. Based on Harness Primitives, we propose STITCH, a framework that Selects suitable primitives given Task Information and compiles them into Task-speCific Harnesses at test time. This separation enables task-specific harnesses without generating or repairing mechanism code at test time. Extensive experiments demonstrate that STITCH not only improves harness adaptability and robustness, but also scales with the primitive library size, boosting task success rates by up to 12 points over fixed harness baselines, surpassing human-designed harnesses like Codex CLI while maintaining a minimal test-time harness composition overhead of only 2.7%, 638 times more efficient than generating task-specific harnesses from scratch. Ultimately, our work demonstrates that building task-adaptive harnesses can be beneficial for completing diverse tasks and that building reusable primitives can be a promising path towards this goal.
Figures & tables
Figure 1: Overview of STITCH. Development stage turns recurring failures into a library of implemented primitives, and test-time harness composition selects and connects these primitives for individual tasks.
Method
Overall
Django
SymPy
Other repos
Pass@1
Pass@ k
Pass^k
Pass@1
Pass@ k
Pass^k
Pass@1
Pass@ k
Pass^k
Pass@1
Pass@ k
Pass^k
Mini-SWE-agent (Seed Harness)
73.0
78.0
68.0
75.0
78.8
71.2
75.0
78.6
71.4
69.1
76.5
61.8
SWE-agent
47.5
63.0
32.0
53.8
71.2
36.5
42.9
50.0
35.7
39.7
55.9
23.5
Codex CLI
79.0
83.0
75.0
78.8
80.8
76.9
64.3
71.4
57.1
85.3
91.2
79.4
MemoHarness
70.5
77.0
64.0
70.2
76.9
63.5
67.9
71.4
64.3
72.1
79.4
64.7
MetaHarness
73.5
79.0
68.0
74.0
78.8
69.2
64.3
71.4
57.1
76.5
82.4
70.6
Table 1: The performance of the GPT-5.6-Luna on SWE-bench Verified.
Method
Overall
Easy
Medium
Hard
Pass@1
Pass@ k
Pass^k
Pass@1
Pass@ k
Pass^k
Pass@1
Pass@ k
Pass^k
Pass@1
Pass@ k
Pass^k
Mini-SWE-agent (Seed Harness)
60.0
68.9
51.1
100.0
100.0
100.0
65.5
75.9
55.2
38.5
46.2
30.8
SWE-agent
13.3
17.8
8.9
33.3
33.3
33.3
13.8
17.2
10.3
7.7
15.4
0.0
Codex CLI
56.7
66.7
46.7
66.7
66.7
66.7
67.2
79.3
55.2
30.8
38.5
23.1
MemoHarness
53.3
66.7
40.0
83.3
100.0
66.7
62.1
75.9
48.3
26.9
38.5
15.4
MetaHarness
55.6
66.7
44.4
66.7
100.0
33.3
62.1
72.4
51.7
38.5
46.2
30.8
Table 2: The performance of the GPT-5.6-Luna on Terminal-Bench 2.
Figure 2: The scalability of STITCH on Harness Primitives Library size. Left : The performance of STITCH on SWE-bench Verified scales with Harness Primitives Library size. Right : The performance of STITCH on Terminal-bench 2 scales with Harness Primitives Library size.
Method
Overall
GPT-5.6-Terra
DeepSeek-V4-Flash
Claude-4.5-Haiku
Pass@1
Pass@ k
Pass k
Pass@1
Pass@ k
Pass k
Pass@1
Pass@ k
Pass k
Pass@1
Pass@ k
Pass k
Mini-SWE-agent (Seed Harness)
75.2
80.3
70.0
82.0
84.0
80.0
82.5
89.0
76.0
61.0
68.0
54.0
SWE-agent
73.3
79.7
67.0
74.5
83.0
66.0
89.5
93.0
86.0
56.0
63.0
49.0
Codex CLI
78.5
84.7
72.3
88.0
92.0
84.0
82.0
90.0
74.0
65.5
72.0
59.0
Our developed primitives (fixed across tasks)
Change-surface Tracer
77.7
83.7
71.7
79.5
84.0
75.0
87.5
92.0
83.0
66.0
75.0
57.0
Table 3: Out-of-development-model actor generalization on SWE-bench Verified.
Figure 3: The cross domain generalization of STITCH. Left : The generalization of STITCH on SWE-bench with out-of-domain primitives developed on Terminal-bench. Right : The generalization of STITCH on Terminal-bench with out-of-domain primitives developed on SWE-bench.
Figure 4: Left : Despite different selection rates of primitives, all primitives are fully executable and activated at test time. Right : The cost overhead of STITCH is as low as 2.7% compared to the actor. In comparison, even the lower bound cost of coding the harness from scratch is 17.23 × higher than the actor cost, making STITCH 638 × efficient.
Appendix figures & tables7 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 5: An illustrative State Guard primitive. The same file-preservation mechanism applies to different task-critical inputs without implementing a domain-specific solver. The case summarizes development records rather than reproducing a verbatim trajectory; it is not a controlled estimate of improvement or evidence of held-out generalization. Protection covers designated local file contents, not arbitrary external side effects.
Figure 6: An illustrative composition intent. The composer names reusable operations and explains their purpose; the deterministic compiler supplies the required dependencies and graph connections. Aliases identify instances within the intent, so the composer does not need to assign runtime node identifiers.
Figure 7: Example harness-composer prompt.
Figure 8: The distribution of the selected primitives by the composer.
Figure 9: The combination of multi-primitive selections by the composer.
Primitive
Application scope
Operation and use of its output
Change-surface Tracer
One public behavior spans coupled implementations or representations that must obey the same rule.
Identifies the behavior owner and an alternate implementation path to inspect. The resulting checks direct attention to related edits or tests that a single-file repair could miss.
Compatibility-envelope Gate
A requested change must preserve a concrete legacy input, entry point, or previously supported behavior.
Pairs the requested behavior with preservation and alternate legacy checks. These checks make the compatibility requirement explicit when the actor chooses and verifies a patch.
Contract-case Explorer
A nearby case or an interaction between features can expose an incomplete interpretation of the requested behavior.
Develops a reported case, a preservation control, and a contrasting alternate case. Their expected observations help distinguish a narrow patch from a repair covering the behavioral boundary.
Discriminating-oracle Runner
An ordinary return-value or exception check could pass while a relevant side effect remains wrong.
Specifies a focused observation of behavior such as mutation, object identity, callback order, or generated structure, together with a preservation control. This gives the actor a more discriminating correctness check.
Domain-trace Template
A value or reference crosses changes in meaning, ownership, or lifecycle state.
Organizes checks around the relevant transitions and an invariant that should survive them. The trace guides the actor toward the point where the value first acquires the wrong meaning or state.
Environment-capability Runner
A missing executable, dependency, backend, locale, or operating-system facility may prevent the reported behavior from being exercised.
Specifies a capability probe alongside behavior and preservation checks. The probe helps the actor distinguish unavailable environment support from an observed defect before interpreting a test result.
Appendix
Table 4: The six SWE-bench primitives introduced in the main paper. Application scopes guide selection; the described checks and evidence guide repository inspection, editing, and verification.
Primitive
Application scope
Operation and use of its output
State Guard
Irreplaceable source files may be modified by task commands, while the required output is stored separately.
Saves source contents and hashes, detects changes, and restores modified files. Reports the affected paths so the actor can work on a copy or a separate output; checks source integrity before completion.
Check Replay
A concrete executable check is available and should remain valid as the actor revises the deliverable.
Retains the check command and its execution context, then reruns the same check at completion. Records the result and restores local workspace writes made by the check, exposing regressions without weakening the acceptance criterion.
First-Failure Localizer
Reference and candidate commands produce comparable ordered observations whose first difference can identify a repair target.
Runs both commands from the same initial workspace state and compares exit status, standard output, and error output. Returns the matching prefix and first difference to focus the actor’s next investigation.
Experiment Keeper
Iterative search produces candidate files that can be compared using a numeric objective with a known optimization direction.
Repeatedly measures valid candidates, compares their median scores, and retains the best candidate’s contents. Restores the retained candidate after a worse trial and remeasures it before completion.
Execution Supervisor
A background job or persistent service needs explicit lifecycle management, including a readiness check when appropriate.
Manages the process group, deadline, logs, and cleanup. Records job completion or probes service readiness, allowing the actor to proceed using an observed execution result rather than repeated manual launch-and-poll steps.
Appendix
Table 5: The five Terminal-Bench primitives introduced in the main paper. Each operation uses inputs available from the task or ordinary actor execution and returns evidence for the next action or completion check.