Composing Task-specific Agent Harnesses at Test Time with Reusable Primitives
Organizations: University of Illinois Urbana-Champaign · University of Michigan, Ann Arbor
Abstract
Agent harnesses govern how large language models (LLMs) gather context, invoke tools, verify results, preserve state, and terminate, largely affecting agent performance. However, the value of each harness mechanism can differ across heterogeneous tasks: a mechanism that improves one task may impose overhead or context distraction on another, leading to the suboptimality of a global harness. We characterize this suboptimality as a mismatch induced by fixed mechanism choices, motivating task-specific harness construction. Nonetheless, generating harness code for each task introduces generation and debugging costs, with execution risks that can compound as more mechanisms are generated. To address those challenges, we introduce Harness Primitives, reusable harness mechanisms with clear application scope and composition contract mined from failed task trajectories. Based on Harness Primitives, we propose STITCH, a framework that Selects suitable primitives given Task Information and compiles them into Task-speCific Harnesses at test time. This separation enables task-specific harnesses without generating or repairing mechanism code at test time. Extensive experiments demonstrate that STITCH not only improves harness adaptability and robustness, but also scales with the primitive library size, boosting task success rates by up to 12 points over fixed harness baselines, surpassing human-designed harnesses like Codex CLI while maintaining a minimal test-time harness composition overhead of only 2.7%, 638 times more efficient than generating task-specific harnesses from scratch. Ultimately, our work demonstrates that building task-adaptive harnesses can be beneficial for completing diverse tasks and that building reusable primitives can be a promising path towards this goal.
Figures & tables
| Method | Overall | Django | SymPy | Other repos | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Pass@1 | Pass@ | Pass^k | Pass@1 | Pass@ | Pass^k | Pass@1 | Pass@ | Pass^k | Pass@1 | Pass@ | Pass^k | |
| Mini-SWE-agent (Seed Harness) | 73.0 | 78.0 | 68.0 | 75.0 | 78.8 | 71.2 | 75.0 | 78.6 | 71.4 | 69.1 | 76.5 | 61.8 |
| SWE-agent | 47.5 | 63.0 | 32.0 | 53.8 | 71.2 | 36.5 | 42.9 | 50.0 | 35.7 | 39.7 | 55.9 | 23.5 |
| Codex CLI | 79.0 | 83.0 | 75.0 | 78.8 | 80.8 | 76.9 | 64.3 | 71.4 | 57.1 | 85.3 | 91.2 | 79.4 |
| MemoHarness | 70.5 | 77.0 | 64.0 | 70.2 | 76.9 | 63.5 | 67.9 | 71.4 | 64.3 | 72.1 | 79.4 | 64.7 |
| MetaHarness | 73.5 | 79.0 | 68.0 | 74.0 | 78.8 | 69.2 | 64.3 | 71.4 | 57.1 | 76.5 | 82.4 | 70.6 |
| Method | Overall | Easy | Medium | Hard | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Pass@1 | Pass@ | Pass^k | Pass@1 | Pass@ | Pass^k | Pass@1 | Pass@ | Pass^k | Pass@1 | Pass@ | Pass^k | |
| Mini-SWE-agent (Seed Harness) | 60.0 | 68.9 | 51.1 | 100.0 | 100.0 | 100.0 | 65.5 | 75.9 | 55.2 | 38.5 | 46.2 | 30.8 |
| SWE-agent | 13.3 | 17.8 | 8.9 | 33.3 | 33.3 | 33.3 | 13.8 | 17.2 | 10.3 | 7.7 | 15.4 | 0.0 |
| Codex CLI | 56.7 | 66.7 | 46.7 | 66.7 | 66.7 | 66.7 | 67.2 | 79.3 | 55.2 | 30.8 | 38.5 | 23.1 |
| MemoHarness | 53.3 | 66.7 | 40.0 | 83.3 | 100.0 | 66.7 | 62.1 | 75.9 | 48.3 | 26.9 | 38.5 | 15.4 |
| MetaHarness | 55.6 | 66.7 | 44.4 | 66.7 | 100.0 | 33.3 | 62.1 | 72.4 | 51.7 | 38.5 | 46.2 | 30.8 |
| Method | Overall | GPT-5.6-Terra | DeepSeek-V4-Flash | Claude-4.5-Haiku | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Pass@1 | Pass@ | Pass k | Pass@1 | Pass@ | Pass k | Pass@1 | Pass@ | Pass k | Pass@1 | Pass@ | Pass k | |
| Mini-SWE-agent (Seed Harness) | 75.2 | 80.3 | 70.0 | 82.0 | 84.0 | 80.0 | 82.5 | 89.0 | 76.0 | 61.0 | 68.0 | 54.0 |
| SWE-agent | 73.3 | 79.7 | 67.0 | 74.5 | 83.0 | 66.0 | 89.5 | 93.0 | 86.0 | 56.0 | 63.0 | 49.0 |
| Codex CLI | 78.5 | 84.7 | 72.3 | 88.0 | 92.0 | 84.0 | 82.0 | 90.0 | 74.0 | 65.5 | 72.0 | 59.0 |
| Our developed primitives (fixed across tasks) | ||||||||||||
| Change-surface Tracer | 77.7 | 83.7 | 71.7 | 79.5 | 84.0 | 75.0 | 87.5 | 92.0 | 83.0 | 66.0 | 75.0 | 57.0 |
Appendix figures & tables7 assets
Supplementary material from the paper’s appendix.
Appendix
| Primitive | Application scope | Operation and use of its output |
|---|---|---|
| Change-surface Tracer | One public behavior spans coupled implementations or representations that must obey the same rule. | Identifies the behavior owner and an alternate implementation path to inspect. The resulting checks direct attention to related edits or tests that a single-file repair could miss. |
| Compatibility-envelope Gate | A requested change must preserve a concrete legacy input, entry point, or previously supported behavior. | Pairs the requested behavior with preservation and alternate legacy checks. These checks make the compatibility requirement explicit when the actor chooses and verifies a patch. |
| Contract-case Explorer | A nearby case or an interaction between features can expose an incomplete interpretation of the requested behavior. | Develops a reported case, a preservation control, and a contrasting alternate case. Their expected observations help distinguish a narrow patch from a repair covering the behavioral boundary. |
| Discriminating-oracle Runner | An ordinary return-value or exception check could pass while a relevant side effect remains wrong. | Specifies a focused observation of behavior such as mutation, object identity, callback order, or generated structure, together with a preservation control. This gives the actor a more discriminating correctness check. |
| Domain-trace Template | A value or reference crosses changes in meaning, ownership, or lifecycle state. | Organizes checks around the relevant transitions and an invariant that should survive them. The trace guides the actor toward the point where the value first acquires the wrong meaning or state. |
| Environment-capability Runner | A missing executable, dependency, backend, locale, or operating-system facility may prevent the reported behavior from being exercised. | Specifies a capability probe alongside behavior and preservation checks. The probe helps the actor distinguish unavailable environment support from an observed defect before interpreting a test result. |
| Primitive | Application scope | Operation and use of its output |
|---|---|---|
| State Guard | Irreplaceable source files may be modified by task commands, while the required output is stored separately. | Saves source contents and hashes, detects changes, and restores modified files. Reports the affected paths so the actor can work on a copy or a separate output; checks source integrity before completion. |
| Check Replay | A concrete executable check is available and should remain valid as the actor revises the deliverable. | Retains the check command and its execution context, then reruns the same check at completion. Records the result and restores local workspace writes made by the check, exposing regressions without weakening the acceptance criterion. |
| First-Failure Localizer | Reference and candidate commands produce comparable ordered observations whose first difference can identify a repair target. | Runs both commands from the same initial workspace state and compares exit status, standard output, and error output. Returns the matching prefix and first difference to focus the actor’s next investigation. |
| Experiment Keeper | Iterative search produces candidate files that can be compared using a numeric objective with a known optimization direction. | Repeatedly measures valid candidates, compares their median scores, and retains the best candidate’s contents. Restores the retained candidate after a worse trial and remeasures it before completion. |
| Execution Supervisor | A background job or persistent service needs explicit lifecycle management, including a readiness check when appropriate. | Manages the process group, deadline, logs, and cleanup. Records job completion or probes service readiness, allowing the actor to proceed using an observed execution result rather than repeated manual launch-and-poll steps. |