Identical Runs, Different Results: Benchmarking AI Coding Agents on Open-Weight Models
Organizations: Central European University · Epoch
Abstract
Repeated runs of the same coding agent are known to give different benchmark scores. We ask what that variation means for a team running an agent on its own task, by intensive replication on one machine-learning task: an agent improves the training code of an XGBoost classifier for airline delays, and a holdout it never sees scores the result. Across 584 runs, we compare six agents on six open-weight model endpoints, run six agent-model pairings 52 times each under fixed settings, and repeat three of them on a larger model from the same family. Identical runs of one pairing varied more than the pairings differed from one another, so comparisons of a few runs ranked them unreliably; resolving the agent differences we observed would take tens to more than a hundred runs of each. Runs on the larger model scored clearly higher, but by less than one run-to-run standard deviation, and the gap was more than twice as large with one agent as with the others. Fewer than one run in twenty broke the task's data rules, but those runs held the highest scores. Rejecting those runs first and keeping the best compliant result among a few attempts reliably improved the delivered model, even though a few runs could not rank the agents. On flights from a later year, the delivered models kept only a third of their gain over the starting code. At list prices, cost differed more than twentyfold between two agents on the same model, mostly through the prompt cache. Agents and models should be evaluated as pairings, over repeated attempts, with compliance reported beside quality. Data, code and every delivered program: https://github.com/earino/identical-runs-different-results
Figures & tables
| Study 1 | Study 2 | Study 3 | |
|---|---|---|---|
| Question | does the pairing matter? | how much does one run vary? | what does a larger model buy? |
| Agents | all six | pi, OpenCode, Hermes | pi, OpenCode, Hermes |
| Models | six endpoints | GLM-5.3 Flash, DeepSeek 4.1 Flash | GLM-5.3 (against Study 2’s GLM-5.3 Flash) |
| Runs | 3 per pairing; 116 in all, 113 delivered an artifact, 103 scored | 52 per pairing; 312, all scored | 52 per pairing; 156, 152 scored |
| Dates | 13 to 15 September 2026 | 15 to 16 September 2026 | 16 to 17 September 2026 |
| Machines | one container per run | four identical 16-core machines, four runs at a time | four more machines of the same type |
| Result | Data | How it arose |
|---|---|---|
| Spread within against between pairings | Study 2 | planned |
| How often few-run comparisons err | Study 2, every possible draw | planned question, analysed after the runs |
| Rule-breaking in the upper tail | Studies 2 and 3, every run | audited throughout, not a planned comparison |
| Repeat-and-select policy | Study 2 | retrospective |
| Model substitution | Study 3 against Study 2’s Flash runs | planned; arms a day apart |
| Agent-by-model interaction | Studies 2 and 3 | exploratory |
| Agent and model | Compliant runs | Mean | Median | SD | 95% interval of the mean | Best | All-run mean |
|---|---|---|---|---|---|---|---|
| pi, GLM-5.3 Flash | 52 | 0.7361 | 0.7359 | 0.0104 | 0.7332 – 0.7389 | 0.7597 | 0.7361 |
| Hermes, GLM-5.3 Flash | 47 | 0.7379 | 0.7404 | 0.0112 | 0.7347 – 0.7411 | 0.7590 | 0.7427 |
| OpenCode, GLM-5.3 Flash | 51 | 0.7400 | 0.7402 | 0.0097 | 0.7373 – 0.7427 | 0.7596 | 0.7417 |
| OpenCode, DeepSeek 4.1 Flash | 52 | 0.7403 | 0.7418 | 0.0096 | 0.7377 – 0.7429 | 0.7578 | 0.7403 |
| Hermes, DeepSeek 4.1 Flash | 49 | 0.7441 | 0.7437 | 0.0109 | 0.7411 – 0.7472 | 0.7663 | 0.7459 |
| pi, DeepSeek 4.1 Flash | 51 | 0.7456 | 0.7462 | 0.0125 | 0.7421 – 0.7490 | 0.7695 | 0.7455 |
| Model row | Agent gap | Run-to-run gap | Ratio |
|---|---|---|---|
| DeepSeek V4 Flash, Ollama | 0.021 | 0.021 | 1.0 |
| DeepSeek V4.1 Flash, Ollama | 0.019 | 0.018 | 1.0 |
| Nemotron Super, Ollama | 0.011 | 0.010 | 1.1 |
| DeepSeek 4.1 Flash, LunaRoute | 0.036 | 0.027 | 1.3 |
| GLM-5.3 Flash, LunaRoute | 0.022 | 0.012 | 1.9 |
| GLM-5.3, LunaRoute | 0.032 | 0.011 | 3.0 |
| Rule | Runs | Spread of pairing means | Median SD | Best | Best of 3, median | Best of 10, median |
|---|---|---|---|---|---|---|
| Exclude nothing | 312 | 0.0098 | 0.0133 | 0.8293 | 0.7499 | 0.7565 |
| Exclude evaluation-label training | 307 | 0.0098 | 0.0111 | 0.8036 | 0.7493 | 0.7548 |
| Exclude batch features | 307 | 0.0095 | 0.0117 | 0.8293 | 0.7498 | 0.7564 |
| Exclude both (the paper’s rule) | 302 | 0.0095 | 0.0107 | 0.7695 | 0.7493 | 0.7548 |
| Attempts | At least one compliant | Median kept | 5th percentile | 95th percentile | Gain over one attempt (95% interval) |
|---|---|---|---|---|---|
| 1 | 96.8% (94.2 to 98.2) | 0.7412 | 0.7220 | 0.7587 | |
| 3 | at least 99.88% | 0.7493 | 0.7355 | 0.7641 | +0.0081 (+0.0063 to +0.0098) |
| 5 | at least 99.99% | 0.7526 | 0.7409 | 0.7645 | +0.0114 (+0.0085 to +0.0131) |
| 10 | at least 99.99% | 0.7548 | 0.7460 | 0.7663 | +0.0136 (+0.0113 to +0.0154) |
| Agent | GLM-5.3 | Runs | GLM-5.3 Flash | Runs | Gain | 95% interval |
|---|---|---|---|---|---|---|
| pi | 0.7513 | 47 | 0.7361 | 52 | +0.0152 | +0.0110 – +0.0193 |
| Hermes | 0.7446 | 47 | 0.7379 | 47 | +0.0067 | +0.0021 – +0.0113 |
| OpenCode | 0.7455 | 50 | 0.7400 | 51 | +0.0055 | +0.0016 – +0.0094 |
| All three | 0.7471 | 144 | 0.7380 | 150 | +0.0091 | +0.0066 – +0.0115 |
| Model | Runs | Compliant | Yield | Compliant mean | One attempt, median | Three attempts, delivered | Three attempts, median |
|---|---|---|---|---|---|---|---|
| GLM-5.3 Flash | 156 | 150 | 96.2% | 0.7380 | 0.7385 | at least 99.76% | 0.7473 |
| GLM-5.3 | 156 | 144 | 92.3% | 0.7471 | 0.7482 | at least 99.67% | 0.7564 |
| Question | Effect | In run-to-run SDs | Runs per arm |
|---|---|---|---|
| Did the agent improve on the starting code? | +0.0277 | 2.7 | — |
| Did the larger model improve the pairing? | +0.0091 | 0.9 | 21 |
| Are two agents different on GLM-5.3? | +0.0067 | 0.6 | 39 |
| Are two agents different on GLM-5.3 Flash? | +0.0039 | 0.4 | 113 |
| Study: model change | Contrast | Estimate | SE | In SEs | Pointwise 95% interval | Holm p |
|---|---|---|---|---|---|---|
| 2: to DeepSeek 4.1 Flash | pi vs Hermes | +0.0033 | 0.0032 | 1.0 | 0.0029 to +0.0094 | 0.30 |
| pi vs OpenCode | +0.0092 | 0.0030 | 3.1 | +0.0034 to +0.0150 | 0.006 | |
| Hermes vs OpenCode | +0.0059 | 0.0030 | 2.0 | +0.0002 to +0.0116 | 0.095 | |
| 3: to GLM-5.3 | pi vs Hermes | +0.0085 | 0.0032 | 2.7 | +0.0024 to +0.0148 | 0.015 |
| pi vs OpenCode | +0.0097 | 0.0029 | 3.3 | +0.0040 to +0.0155 | 0.003 | |
| Hermes vs OpenCode | +0.0012 | 0.0031 | 0.4 | 0.0049 to +0.0072 | 0.71 |
| 2006 holdout | 2007 flights | |
|---|---|---|
| Gain over the starting code, 446 compliant runs | +0.0280 | +0.0092 |
| Compliant runs below the starting code | 0 | 33 |
| Spread of the six Study 2 pairing means | 0.0095 | 0.0021 |
| Median run-to-run SD, Study 2 | 0.0107 | 0.0060 |
| Best of three attempts over one | +0.0081 | +0.0022 |
| Larger model’s gain | +0.0091 (7.1 SE) | +0.0027 (3.8 SE) |
| Agent | Run 1 | Run 2 | Run 3 | Mean cost |
|---|---|---|---|---|
| pi | 98%, $0.08 | 98%, $0.11 | 97%, $0.05 | $0.08 |
| OpenClaw | no record | 98%, $0.11 | 97%, $0.07 | $0.09 |
| Codex | 96%, $0.14 | 96%, $0.19 | 92%, $0.13 | $0.15 |
| OpenCode | 80%, $0.24 | 64%, $0.19 | 66%, $0.49 | $0.31 |
| Hermes | 56%, $1.09 | 89%, $0.47 | 74%, $0.24 | $0.60 |
| Claude Code | 11%, $1.41 | 9%, $2.95 | 2%, $1.21 | $1.86 |
| Agent, model | Input tokens per run | Output tokens per run | List-price cost per run | Yield | Compliant mean | Best of 3, median |
|---|---|---|---|---|---|---|
| pi, GLM-5.3 Flash | 4.34M | 58k | $0.45 | 100% | 0.7361 | 0.7460 |
| Hermes, GLM-5.3 Flash | 8.12M | 68k | $0.83 | 90.4% | 0.7379 | 0.7461 |
| OpenCode, GLM-5.3 Flash | 5.70M | 32k | $0.58 | 98.1% | 0.7400 | 0.7486 |
| pi, DeepSeek 4.1 Flash | 3.93M | 63k | $0.63 | 98.1% | 0.7456 | 0.7567 |
| Hermes, DeepSeek 4.1 Flash | 8.65M | 69k | $1.34 | 94.2% | 0.7441 | 0.7535 |
| OpenCode, DeepSeek 4.1 Flash | 5.07M | 26k | $0.78 | 100% | 0.7403 | 0.7488 |
| Cases reviewed each month | Caught by the 0.7695 model | Caught by the 0.7471 model | Difference |
|---|---|---|---|
| 5,000 | 562 | 603 | $4,501 |
| 20,000 | 1,351 | 1,287 | $7,031 |
| 100,000 | 3,714 | 3,462 | $27,665 |
Appendix figures & tables4 assets
Supplementary material from the paper’s appendix.
Appendix
| Model rows | Scheduled | Started | Delivered an artifact | Compliant |
|---|---|---|---|---|
| Five complete rows | 90 | 90 | 90 | 88 |
| DeepSeek 4.1 Flash, LunaRoute | 18 | 15 | 15 | 15 |
| DeepSeek V4 Pro, Ollama (partial) | 8 | 8 | 8 | not assessed |
| Subset | Runs | Rank correlation | 95% interval | p |
|---|---|---|---|---|
| All runs | 312 | +0.55 | +0.45 to +0.63 | below 0.001 |
| Compliant runs | 302 | +0.55 | +0.46 to +0.63 | below 0.001 |
| Compliant, above a quarter of budget | 209 | +0.43 | +0.29 to +0.54 | below 0.001 |
| Compliant, above half | 110 | +0.29 | +0.10 to +0.45 | 0.001 |
| Compliant, above three quarters | 70 | +0.12 | 0.15 to +0.41 | 0.342 |
| Pairing | 1 attempt | 3 | 5 | 10 | Gain, 1 to 3 | Gain, 1 to 10 |
|---|---|---|---|---|---|---|
| pi, GLM-5.3 Flash | 0.7359 | 0.7460 | 0.7493 | 0.7529 | +0.0101 | +0.0170 |
| Hermes, GLM-5.3 Flash | 0.7404 | 0.7461 | 0.7499 | 0.7543 | +0.0057 | +0.0139 |
| OpenCode, GLM-5.3 Flash | 0.7402 | 0.7486 | 0.7518 | 0.7532 | +0.0084 | +0.0130 |
| OpenCode, DeepSeek 4.1 Flash | 0.7418 | 0.7488 | 0.7506 | 0.7533 | +0.0070 | +0.0115 |
| Hermes, DeepSeek 4.1 Flash | 0.7437 | 0.7535 | 0.7550 | 0.7607 | +0.0098 | +0.0170 |
| pi, DeepSeek 4.1 Flash | 0.7462 | 0.7567 | 0.7590 | 0.7613 | +0.0105 | +0.0151 |
| Model | Price per M tokens, in / out | Repriced ledger | If 90 percent of input were cached |
|---|---|---|---|
| GLM-5.3 Flash and DeepSeek 4.1 Flash, as run | 0.33 and 0.60 | $244 | $54 |
| GLM-5.3 | 4.40 | $2,718 | $807 |
| Gemini 3.1 Pro | 12.00 | $4,026 | $1,009 |
| Kimi K3 | 13.28 | $5,268 | $1,335 |
| Claude Opus 5 | 25.00 | $9,939 | $2,397 |
| GPT-5.5 Pro | 180.00 | $60,385 | N/A |