cs.AISep 30, 2026

How Much of a Harness Does a Strong Agent Need for Autonomous ML Engineering?

Authors: Kirill Brilliantov, Alejandro Hernández-Cano, Emmanuel Abbé

Organizations: EPFL · Apple

Abstract

Recent autonomous machine learning engineering (MLE) agents have made significant progress on public leaderboards. Often motivated by progress stagnation over long-horizon cycles and limited Large Language Model (LLM) primitives, modern MLE agents are deployed on top of increasingly elaborate machinery: multi-agent orchestrators, dedicated retrieval subagents, and more. While such harnesses expand, the use of more primitive but improved coding agents - where LLMs have direct access to the execution environment through read, write, and bash primitives - has received little attention in the field. In this paper we find that, under an equal time budget and the same frontier LLM backbone, open-source state-of-the-art harnesses provide no advantages over a single session of a minimal-harness coding agent baseline, pointing to the backbone as the primary driver for performance. Via a series of large-scale systematic ablation studies, we argue that the machinery layers become redundant in the coding agent setting. We conclude that the effort spent elaborating hand-crafted harnesses around strong models yields poor returns for current MLE benchmarks.

Figures & tables

Appendix figures & tables28 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Agentic Harness Engineering: Observability-Driven Automatic Evolution of Coding-Agent Harnesses

    Apr 28, 2026Jiahang Lin, Shichun Liu, Chengjun Pan +8Agent HarnessCoding Agents

  2. Code as Agent Harness

    May 18, 2026Xuying Ning, Katherine Tieu, Dongqi Fu +39Agent HarnessCoding Agents

  3. An Empirical Study of Harness Design for Coding Agents

    Sep 17, 2026Run-Ze Fan, Zihao Zhang, Simin Ma +6Coding AgentsAgent Harness