How Much of a Harness Does a Strong Agent Need for Autonomous ML Engineering?
Organizations: EPFL · Apple
Abstract
Recent autonomous machine learning engineering (MLE) agents have made significant progress on public leaderboards. Often motivated by progress stagnation over long-horizon cycles and limited Large Language Model (LLM) primitives, modern MLE agents are deployed on top of increasingly elaborate machinery: multi-agent orchestrators, dedicated retrieval subagents, and more. While such harnesses expand, the use of more primitive but improved coding agents - where LLMs have direct access to the execution environment through read, write, and bash primitives - has received little attention in the field. In this paper we find that, under an equal time budget and the same frontier LLM backbone, open-source state-of-the-art harnesses provide no advantages over a single session of a minimal-harness coding agent baseline, pointing to the backbone as the primary driver for performance. Via a series of large-scale systematic ablation studies, we argue that the machinery layers become redundant in the coding agent setting. We conclude that the effort spent elaborating hand-crafted harnesses around strong models yields poor returns for current MLE benchmarks.
Figures & tables
| Percentile ( ) | Medal rate (%, ) | |||
| Self-select | Oracle | Self-select | Oracle | |
| base | 66.51 [62.46, 69.97] | 69.29 [65.82, 72.13] | 55.7 [48.2, 62.9] | 60.4 [53.2, 66.8] |
| +D | 63.62 [59.27, 68.26] | 67.40 [63.22, 71.98] | 45.2 [38.1, 54.8] | 47.0 [38.1, 56.0] |
| +P3 | 67.77 [64.58, 70.62] | 71.79 [69.34, 73.83] | 52.4 [42.8, 61.9] | 60.7 [53.6, 67.3] |
| +B+P3 | 58.91 [53.50, 64.05] | 68.57 [65.77, 71.14] | 33.3 [23.8, 40.5] | 53.0 [45.8, 60.1] |
| Surpassed SOTA (%) | |||
|---|---|---|---|
| Backbone | Malena | AiScientist | MLEvolve |
| GLM-5.2 | 21.7 [15.0, 27.5] | 17.5 [12.5, 22.5] | 10.8 [7.5, 15.0] |
| Kimi-K3 | 26.7 [20.0, 32.5] | 27.1 [20.0, 35.0] | 11.7 [7.5, 15.0] |
Appendix figures & tables28 assets
Supplementary material from the paper’s appendix.
Appendix
| Full competition ID | Split | |
|---|---|---|
| alaska2* † ‖ | alaska2-image-steganalysis | Medium |
| aptos2019 | aptos2019-blindness-detection | Lite |
| billion-word* | billion-word-imputation | Medium |
| cassava* † | cassava-leaf-disease-classification | Medium |
| champs* † ‡ | champs-scalar-coupling | Medium |
| freesound* † | freesound-audio-tagging-2019 | Medium |
| s41467-025-65557-7 | s41592-025-02662-x | s42256-023-00630-8 | s42256-025-01010-0 |
| s41551-024-01257-9 | s41592-025-02826-9 | s42256-023-00636-2 | s42256-025-01026-6 |
| s41551-024-01312-5 | s41592-025-02886-x | s42256-023-00654-0 | s43588-024-00689-2 |
| s41587-024-02414-w | s42256-022-00447-x | s42256-023-00712-7 | s43588-024-00698-1 |
| s41592-023-01940-w | s42256-022-00459-7 | s42256-024-00790-1 | s43588-024-00716-2 |
| s41592-023-02035-2 | s42256-022-00518-z | s42256-024-00795-w | s43588-024-00732-2 |
| s41592-023-02124-2 | s42256-022-00526-z | s42256-024-00815-9 | s43588-024-00757-7 |
| Percentile ( ) | Medal rate (%, ) | ||
| won A/B/= | |||
| Oneshot search strategy | |||
| Best-of-N vs UCB1 | -2.71 [-6.60, +1.44] | -5.3 [-11.8, +1.1] | 4/7/18 |
| Best-of-N vs Greedy | -1.98 [-5.86, +1.89] | -5.3 [-11.2, +0.6] | 3/7/19 |
| Best-of-N vs Chain | -0.51 [-4.66, +3.27] | -4.2 [-11.2, +2.9] | 4/7/18 |
| UCB1 vs Greedy | +0.73 [-2.85, +4.35] | +0.0 [-5.7, +5.7] | 5/5/19 |
| Percentile ( ) | Selection gap ( ) | ||
|---|---|---|---|
| Self-select | Oracle | ||
| Malena | 69.30 [66.74, 71.78] | 71.18 [69.02, 73.46] | 1.876 [0.805, 3.205] |
| Best-of-N | 64.21 [61.19, 67.07] | 71.04 [69.49, 72.66] | 6.832 [4.336, 9.634] |
| UCB1 | 66.92 [64.30, 69.36] | 70.89 [68.45, 73.35] | 3.973 [2.484, 5.367] |
| Greedy | 66.19 [63.43, 68.71] | 68.98 [67.32, 70.56] | 2.788 [1.011, 5.054] |
| Chain | 64.72 [61.97, 67.42] | 66.87 [64.08, 69.57] | 2.152 [1.409, 2.899] |
| Percentile ( ) | Selection gap ( ) | ||
| Self-select | Oracle | ||
| BI | 61.24 [58.85, 63.82] | 63.26 [61.28, 65.38] | 2.024 [1.125, 2.901] |
| base | 62.56 [60.34, 64.63] | 67.99 [66.74, 69.26] | 5.433 [3.681, 7.359] |
| +V | 63.74 [59.93, 67.04] | 67.57 [66.23, 68.69] | 3.833 [0.936, 7.263] |
| HV | 61.55 [59.73, 64.36] | 66.03 [63.38, 68.10] | 4.479 [2.412, 5.944] |
| Percentile ( ) | Medal rate (%, ) | ||
|---|---|---|---|
| won A/B/= | |||
| Gemma 4 31B | -3.97 [-7.98, -0.35]* | +7.9 [+4.3, +11.8]* | 4/1/24 |
| DeepSeek V4 Flash | +1.15 [-5.42, +7.77] | +3.3 [-6.3, +12.4] | 9/6/14 |
| DeepSeek V4 Pro | +9.39 [+3.36, +15.98]* | +12.8 [+2.9, +22.4]* | 11/4/14 |
| GLM 5.2 | +5.52 [+1.93, +9.57]* | +12.0 [+5.7, +18.4]* | 12/2/15 |
| Kimi K3 | +4.51 [+0.83, +8.96]* | +10.5 [+4.3, +18.1]* | 8/4/17 |
| Percentile ( ) | Medal rate (%, ) | |||
| Self-select | Oracle | Self-select | Oracle | |
| Infrastructure ablation (24h, 29 tasks) | ||||
| base | 69.30 [66.74, 71.78] | 71.18 [69.02, 73.46] | 59.4 [54.8, 63.9] | 62.5 [58.4, 66.3] |
| +N | 63.85 [59.73, 67.77] | 65.98 [62.05, 69.83] | 48.9 [43.1, 54.0] | 52.3 [46.6, 57.5] |
| -J | 70.47 [66.89, 74.25] | 72.84 [69.55, 76.10] | 62.1 [56.9, 67.2] | 62.1 [56.9, 67.2] |
| -S-J | 65.33 [62.05, 68.61] | 68.36 [65.16, 71.45] | 50.6 [44.8, 56.3] | 56.9 [50.6, 62.6] |
| Percentile ( ) | Medal rate (%, ) | ||
| won A/B/= | |||
| Infrastructure ablation (24h, 29 tasks) | |||
| base vs +N | +5.46 [+0.70, +10.67]* | +10.6 [+3.7, +17.9]* | 8/1/20 |
| base vs -J | -1.17 [-5.64, +3.34] | -2.6 [-10.2, +4.7] | 3/6/20 |
| base vs -S-J | +3.97 [-0.07, +7.97] | +8.9 [+1.7, +15.9]* | 9/3/17 |
| Orchestration ablation (24h, 14 tasks) | |||
| Task | alaska2 | cassava | champs | freesound | hubmap | imet | kuzushiji | multi-modal | nfl | petfinder |
|---|---|---|---|---|---|---|---|---|---|---|
| Avg | – | 1.2 | – | 1.5 | – | 2.0 | 4.2 | 12.5 | 2.5 | 1.7 |
| Self-selected | Oracle | |||
|---|---|---|---|---|
| Group | Any-medal | Mean percentile | Any-medal | Mean percentile |
| Malena 1M | 57.5% [42.2, 71.5] | 77.1 [68.0, 86.2] | 65.0% [49.5, 77.9] | 78.8 [69.8, 87.8] |
| Malena 64k | 50.0% [35.2, 64.8] | 73.5 [65.3, 81.7] | 50.0% [35.2, 64.8] | 74.5 [66.3, 82.7] |
| Percentile ( ) | Medal rate ( ) | |
| base | 55.15 [54.31, 55.98] | 36.8 [35.5, 38.2] |
| +S+J | 56.58 [54.78, 58.41] | 37.9 [34.5, 41.4] |
| +S+J+N | 55.99 [54.18, 57.85] | 37.1 [33.6, 40.1] |
| +H | 57.08 [55.39, 58.80] | 38.8 [35.8, 41.8] |
| +D | 56.95 [55.12, 58.81] | 39.2 [35.8, 42.7] |
| +V | 54.47 [52.62, 56.38] | 35.9 [32.4, 39.4] |
| Percentile ( ) | Medal rate (%, ) | ||
| won A/B/= | |||
| GLM 5.2 | |||
| base vs +S+J | -1.43 [-3.29, +0.57] | -1.1 [-4.7, +2.7] | 9/10/10 |
| base vs +S+J+N | -0.85 [-2.88, +1.22] | -0.2 [-3.9, +3.2] | 9/10/10 |
| base vs +H | -1.93 [-3.85, -2.9e-03]* | -2.0 [-5.4, +1.3] | 7/10/12 |
| base vs +D | -1.80 [-3.82, +0.23] | -2.4 [-6.0, +1.2] | 9/10/10 |
| Percentile ( ) | Medal rate ( ) | |||
| Self-select | Oracle | Self-select | Oracle | |
| Kimi K3 | ||||
| Malena | 72.75 [70.75, 74.82] | 75.54 [73.90, 77.35] | 60.0 [55.6, 64.4] | 65.6 [60.0, 70.0] |
| Arbor | 68.46 [66.63, 70.08] | 69.53 [67.76, 71.25] | 60.0 [56.7, 63.3] | 61.1 [57.8, 64.4] |
| AiScientist | 64.15 [61.26, 67.10] | 65.48 [62.54, 68.32] | 52.2 [46.7, 57.8] | 53.3 [47.8, 58.9] |
| MLEvolve | 63.89 [62.00, 65.74] | 70.32 [68.98, 71.58] | 50.0 [46.7, 54.4] | 63.3 [60.0, 66.7] |
| Percentile ( ) | Medal rate (%, ) | ||
| won A/B/= | |||
| Kimi K3 | |||
| Malena vs Arbor | +4.29 [+1.54, +7.05]* | +0.0 [-5.6, +5.6] | 5/5/20 |
| Malena vs AiScientist | +8.60 [+4.91, +12.15]* | +7.8 [+0.0, +15.6] | 7/2/21 |
| Malena vs MLEvolve | +8.85 [+6.11, +11.70]* | +10.0 [+3.3, +16.7]* | 7/2/21 |
| GLM 5.2 | |||
| reasoning_effort | Any-medal rate (%) [95% CI] | None-rate (%) | Mean percentile [95% CI] | Median reasoning (chars) | Mean nodes/run | Buggy rate (%) |
|---|---|---|---|---|---|---|
| max | 0.0 [0.0, 0.0] | 13.8 | 42.6 [36.0, 48.5] | 37124 | 18.5 | 61 |
| medium (canonical) | 25.0 [10.0, 30.0] | 2.5 | 53.1 [44.9, 59.8] | 7069 | 62.8 | 62 |
| Configuration | Any-medal rate (%) [95% CI] | Gold (%) | Silver (%) | Bronze (%) | Mean percentile [95% CI] |
|---|---|---|---|---|---|
| resource_smart_llm (canonical) | 28.3 [10.0, 50.0] | 10.0 | 13.3 | 5.0 | 66.6 [52.9, 80.7] |
| resource_smart_policy | 20.0 [0.0, 40.0] | 2.5 | 10.0 | 7.5 | 52.2 [39.6, 65.5] |
| Ablation | Baseline any-medal rate (%) [95% CI] | Ablation any-medal rate (%) [95% CI] |
|---|---|---|
| Executor timeout: 4h 12h | 12.5 [0.0, 50.0] | 12.5 [0.0, 50.0] |
| Convergence early-stop: stop_after 8 999 | 50.0 [0.0, 100.0] | 43.8 [25.0, 75.0] |
| Malena | AiScientist | MLEvolve | |
|---|---|---|---|
| p25 | 8 | 7 | 2 |
| p50 (median) | 16 | 18 | 8 |
| p75 | 36 | 96 | 18 |
| metric | group | p50 | p90 | p95 | p99 |
|---|---|---|---|---|---|
| Dolos containment | real-vs-real | 0.008 | 0.077 | 0.143 | 0.453 |
| Dolos containment | Malena-vs-real | 0.002 | 0.009 | 0.013 | 0.020 |
| mean-pooling cosine | real-vs-real | 0.735 | 0.877 | 0.909 | 0.968 |
| mean-pooling cosine | Malena-vs-real | 0.773 | 0.880 | 0.896 | 0.919 |
| Chamfer | real-vs-real | 0.634 | 0.741 | 0.778 | 0.882 |
| Chamfer | Malena-vs-real | 0.638 | 0.713 | 0.726 | 0.746 |