Harness Engineering for Software Engineering via Modular Executable Dev-Primitives
Organizations: University of Illinois Urbana-Champaign, USA · The Pennsylvania State University, USA
Abstract
Large language models (LLMs) equipped with terminal access have demonstrated strong capabilities in automating software engineering tasks. However, existing agents remain brittle on long-horizon workflows, where they must repeatedly reconstruct program state scattered across source files, configurations, tests, dependencies, and runtime behavior, leading to increasingly long interaction histories, context explosion, and semantic drift. Large repositories further complicate the identification of task-relevant components. To address these challenges, we introduce \textbf{Dev-Primitives} (\emph{Development Primitives}), a modular and executable abstraction that transforms repository components from passive software artifacts into active participants in software engineering. Each Dev-Primitive pairs a repository artifact with a resident LLM, which gives the artifact an agent-native interface grounded in its own implementation and dependencies, enabling natural-language reasoning, inter-component communication, and localized self-modification. Building on Dev-Primitives, we propose \textbf{HERMES}, a Harness Engineering framework for software engineeRing via Modular Executable Dev-PrimitiveS, which instantiates these primitives at repository scale through a dependency-aware dynamic activation mechanism and a bug diagnosis mechanism that maps execution evidence back to the components that must be revised. Extensive experiments on four software engineering benchmarks demonstrate that HERMES outperforms matched baseline harnesses by 12.4% on average. Moreover, when paired with strong activation and diagnosis models, HERMES, even with Qwen3-8B Dev-Primitives, remains within 4.5% of the homogeneous GPT-5.6 Sol configuration across all four benchmarks, while reducing inference cost by 26.2% on Terminal-Bench 4.0, highlighting the importance of harness design in software engineering agents.
Figures & tables
| Baseline Results | HERMES Results | ||||||
| Backbone | Resolved (%) | Cost/Test (\downarrow$ | Latency (s) | Resolved (%) | Cost/Test (\downarrow$ | Latency (s) | Res. (p.p.) |
| Harness | mini-SWE-agent | HERMES | |||||
| GPT-5.6 Sol | 96.20 0.86 | 1.15 | 182.37 | 97.00 0.76 | 1.80 | 205.34 | +0.80 |
| GPT-5.6 Terra | 95.40 0.94 | 0.40 | 180.00 | 96.20 0.86 | 0.98 | 196.71 | +0.80 |
| GPT-5.6 Luna | 93.00 1.14 | 0.04 | 201.07 | 95.60 0.92 | 0.10 | 218.46 | +2.60 |
| Claude Fable 5 | 95.00 0.98 | 2.05 | 356.18 | 96.00 0.88 | 4.50 | 381.27 | +1.00 |
| Baseline Results | HERMES Results | |||||
| Model | Effort | Comp. (%) | Cost (\downarrow$ | Comp. (%) | Cost (\downarrow$ | Comp. (p.p.) |
| Harness | Claude Code | HERMES | ||||
| Claude Opus 5 | medium | 28.5 | 38.4 | 36.5 | 58.6 | +8.0 |
| high | 34.5 | 55.7 | 41.5 | 82.4 | +7.0 | |
| Claude Sonnet 5 | medium | 15.0 | 11.9 | 22.0 | 22.7 | +7.0 |
| high | 6.0 | 24.6 | 25.5 | 34.8 | +19.5 | |
| Baseline Results | HERMES Results | |||||||
| Model | Effort | Resolution (%) | Tokens | Cost (\downarrow$ | Resolution (%) | Tokens | Cost (\downarrow$ | Res. (p.p.) |
| Harness | Codex | HERMES | ||||||
| GPT-5.6 Sol | medium | 33.0 3.4 | 3.6B | 1.9k | 51.8 2.2 | 6.8B | 3.4k | +18.8 |
| max | 37.3 3.8 | 4.4B | 2.5k | 55.5 3.1 | 8.9B | 4.7k | +18.2 | |
| GPT-5.6 Terra | medium | 18.2 3.0 | 4.4B | 1.2k | 44.5 3.3 | 7.2B | 2.2k | +26.3 |
| max | 21.5 3.3 | 5.7B | 1.7k | 48.2 2.5 | 9.1B | 2.9k | +26.7 | |
| Baseline Results | HERMES Results | ||||||||||
| Model | Build | Monitor. | Issue | Test | Avg. | Build | Monitor. | Issue | Test | Avg. | Avg. (p.p.) |
| Harness | Codex | HERMES | |||||||||
| GPT-5.6 Sol | 77.78 | 38.24 | 40.97 | 42.58 | 49.89 | 83.33 | 44.12 | 46.13 | 48.06 | 55.41 | +5.52 |
| GPT-5.6 Terra | 74.07 | 35.29 | 38.71 | 40.32 | 47.10 | 79.63 | 41.18 | 43.55 | 44.84 | 52.30 | +5.20 |
| Harness | Claude Code | HERMES | |||||||||
| Claude Opus 5 | 79.63 | 41.18 | 42.90 | 44.52 | 52.06 | 85.19 | 47.06 | 47.10 | 49.03 | 57.10 | +5.04 |
| Benchmark | Method | Score (%) |
| SWE-bench | MAGIS | 13.9 |
| HERMES | 36.6 | |
| NL2Repo-Bench | CodeTeam (PE) | 34.6 |
| CodeTeam (SFT) | 42.3 | |
| HERMES | 44.9 |
| SWE-bench | SWE Refactor | Terminal-Bench | DevOps-Gym | |||
| Activation | Dev-Primitives | Diagnosis | Resolved (%) | Composite (%) | Resolved (%) | Avg. (%) |
| Homogeneous Backbones | ||||||
| Qwen3-8B | Qwen3-8B | Qwen3-8B | 80.6 | 10.5 | 27.6 | 37.48 |
| GPT-5.6 Luna | GPT-5.6 Luna | GPT-5.6 Luna | 95.6 | 16.0 | 29.1 | 46.64 |
| DeepSeek-V4 | DeepSeek-V4 | DeepSeek-V4 | 82.8 | 15.0 | 39.4 | 53.08 |
| GPT-5.6 Sol | GPT-5.6 Sol | GPT-5.6 Sol | 97.0 | 31.0 | 51.8 | 55.41 |
| Configuration | Resolved (%) | Cost ($) |
| Qwen3-8B / Qwen3-8B / Qwen3-8B | 27.6 | 0.70k |
| Qwen3-8B / Qwen3-8B / GPT-5.5 | 46.1 | 1.56k |
| GPT-5.6 Sol / Qwen3-8B / GPT-5.5 | 49.7 | 2.13k |
| GPT-5.6 Sol / Qwen3-8B / GPT-5.6 Sol | 50.6 | 2.51k |
| GPT-5.6 Sol / GPT-5.6 Sol / GPT-5.6 Sol | 51.8 | 3.40k |
| Benchmark | Recall | Precision | Avg. Act. | Solved R. | Failed R. |
| SWE-bench Verified | 95.4 | 72.6 | 4.7 | 96.1 | 71.4 |
| SWE Refactor Bench | 83.3 | 68.4 | 12.6 | 94.7 | 78.2 |
| Terminal-Bench 4.0 | 84.5 | 64.9 | 6.8 | 93.5 | 74.8 |
| DevOps-Gym | 85.7 | 70.3 | 5.4 | 95.0 | 76.6 |
| Setting | SWE-bench | Refactor | Terminal | DevOps |
| HERMES | 97.0 | 31.0 | 51.8 | 55.41 |
| w/o Inter-Primitive Comm. | 94.2 ( 2.8) | 23.5 ( 7.5) | 46.1 ( 5.7) | 50.34 ( 5.07) |
| w/o On-Demand Activation | 95.8 ( 1.2) | 28.0 ( 3.0) | 49.4 ( 2.4) | 53.26 ( 2.15) |
| w/o Execution Feedback | 93.8 ( 3.2) | 21.5 ( 9.5) | 43.6 ( 8.2) | 48.17 ( 7.24) |
| w/o Diagnosis Feedback | 92.6 ( 4.4) | 20.0 ( 11.0) | 41.2 ( 10.6) | 46.73 ( 8.68) |
Appendix figures & tables8 assets
Supplementary material from the paper’s appendix.
Appendix
| Model | SWE-bench | Refactor | T-Bench | DevOps | Scaling | Ablation |
| OpenAI | ||||||
| GPT-5.6 Sol | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ |
| GPT-5.6 Terra | ✓ | ✓ | ✓ | ✓ | ||
| GPT-5.6 Luna | ✓ | ✓ | ✓ | ✓ | ✓ | |
| GPT-5.5 | ✓ | ✓ | ||||
| GPT-5.4, GPT-5.4 Mini, GPT-5 Mini | ✓ | |||||
| SWE-bench | SWE Refactor | Terminal-Bench | DevOps-Gym | ||
| Backbone | Setting | Resolved (%) | Composite (%) | Resolved (%) | Avg. (%) |
| GPT-5.6 Sol | |||||
| GPT-5.6 Sol | HERMES | 97.0 | 31.0 | 51.8 | 55.41 |
| GPT-5.6 Sol | w/o Inter-Primitive Communication | 94.2 ( 2.8) | 23.5 ( 7.5) | 46.1 ( 5.7) | 50.34 ( 5.07) |
| GPT-5.6 Sol | w/o On-Demand Activation | 95.8 ( 1.2) | 28.0 ( 3.0) | 49.4 ( 2.4) | 53.26 ( 2.15) |
| GPT-5.6 Sol | w/o Execution Feedback | 93.8 ( 3.2) | 21.5 ( 9.5) | 43.6 ( 8.2) | 48.17 ( 7.24) |
| SWE-bench | SWE Refactor | Terminal-Bench | DevOps-Gym | |
| 0 | 92.2 ( 4.8) | 18.5 ( 12.5) | 39.7 ( 12.1) | 45.91 ( 9.50) |
| 1 | 94.6 ( 2.4) | 24.0 ( 7.0) | 45.2 ( 6.6) | 50.17 ( 5.24) |
| 2 | 96.4 ( 0.6) | 28.5 ( 2.5) | 48.8 ( 3.0) | 53.72 ( 1.69) |
| 3 | 97.0 | 31.0 | 51.8 | 55.41 |
| 4 | 97.2 ( 0.2) | 31.5 ( 0.5) | 52.4 ( 0.6) | 55.79 ( 0.38) |
| 5 | 97.2 ( 0.2) | 31.5 ( 0.5) | 52.7 ( 0.9) | 55.88 ( 0.47) |
| SWE-bench Verified | SWE Refactor Bench | Terminal-Bench 4.0 | DevOps-Gym | |||||||||
| Setting | Res. | Tok. | Cost | Comp. | Tok. | Cost | Res. | Tok. | Cost | Avg. | Tok. | Cost |
| HERMES | 97.0 | 1.00 | 1.00 | 31.0 | 1.00 | 1.00 | 51.8 | 1.00 | 1.00 | 55.41 | 1.00 | 1.00 |
| Single editor | 93.4 | 0.86 | 0.85 | 16.8 | 0.90 | 0.89 | 38.8 | 0.86 | 0.84 | 44.21 | 0.87 | 0.86 |
| 3.6 | 14.2 | 13.0 | 11.20 | |||||||||
| Setting | Resolution (%) | Tokens | Cost ($) |
| GPT-5.6 Sol | |||
| Single editor | 38.8 | 5.9B | 2.90k |
| Single editor + compute matching | 45.5 | 6.7B | 3.35k |
| HERMES | 51.8 | 6.8B | 3.40k |
| Avg. Primitive Calls | |||
| Benchmark | HERMES | w/o Activation | Relative Cost w/o Activation |
| SWE-bench Verified | 9.6 | 28.7 | 1.61 |
| SWE Refactor Bench | 25.1 | 64.3 | 1.88 |
| Terminal-Bench 4.0 | 14.2 | 39.6 | 1.73 |
| DevOps-Gym | 11.7 | 33.2 | 1.67 |
| Component | Action | Initial local objective |
| db/models/fields/json.py | modify | Preserve non-ASCII characters when preparing JSONField values |
| forms/fields.py | modify | Update prepare_value to preserve non-ASCII characters |
| forms/utils.py | modify | Avoid ASCII escaping in JSON-formatted form output |
| core/serializers/json.py | modify | Preserve non-ASCII characters in core JSON serialization |
| forms/widgets.py | inspect | Inspect the rendering path and communicate relevant constraints |
| utils/encoding.py | inspect | Inspect shared encoding behavior and communicate Unicode-related constraints |
| From | To | Message (abridged) | Local consequence |
| forms/widgets.py | forms/fields.py | Use ensure_ascii=False when preparing JSON values so Unicode is preserved during rendering. | Refine prepare_value |
| db/models/fields/json.py | core/serializers/json.py | Keep JSON encoding behavior consistent across model-field and serializer paths. | Refine serializer behavior |
| utils/encoding.py | core/serializers/json.py | Preserve Unicode rather than emitting ASCII escape sequences during serialization. | Reinforce serializer objective |