Although effective, coding agents often incur substantial monetary costs. Their recurring cost-inefficient behaviors remain underexplored. We conduct the first study of behavioral cost inefficiencies in coding agents, analyzing 1,200 trajectories from Claude Code and Mini-SWE-Agent across four configurations on SWE-bench Verified. We identify three cost-inefficient behaviors: subsumed retrieval, similar script generation, and test re-execution. We then evaluate three mitigation strategies: structure-aware retrieval, agent-synthesized skills, and developer-designed skills, over 10k trajectories on held-out SWE-bench Verified and Pro tasks. Our main findings are: (1) The three behaviors affect 79.00%--98.00% of coding tasks and account for up to 22.75% of task cost. (2) Structure-aware retrieval can introduce retrieval overhead and alter agent delegation, causing inconsistent improvements in retrieval efficiency and cost increases of up to 28.14%. (3) Agent-synthesized skills tend to produce low-level, trace-specific guidance, limiting their effectiveness and generality. (4) In contrast, developer-designed skills provide high-level, trace-agnostic guidance, reducing cost by up to 41.73%, roughly twice the maximum gain from agent-synthesized skills.
Figures & tables
Figure 1: Cost-inefficient behaviors of Claude Code on SWE-bench Verified task Django -13158 .
CC
MSA S46
MSA MM3
MSA Q35+
Behaviors
Tasks
Freq.
Cost
Tasks
Freq.
Cost
Tasks
Freq.
Cost
Tasks
Freq.
Cost
SubRetrv
64.33%
2.15
5.01%
87.33%
3.45
8.42%
92.33%
5.14
7.88%
89.33%
6.37
11.41%
SimScrpt
20.67%
0.43
1.02%
51.33%
2.54
9.57%
68.00%
4.29
8.85%
57.67%
3.20
7.91%
ReTest
49.67%
2.09
0.83%
69.00%
2.51
3.17%
83.00%
5.29
5.39%
66.00%
2.10
3.43%
Total
79.00%
4.67
6.86%
97.33%
8.50
21.16%
98.00%
14.72
22.12%
96.67%
11.68
22.75%
Table 1: Prevalence, frequency, and cost contribution of cost-inefficient behaviors across agent configurations. Tasks : percentage of tasks exhibiting the behavior; Freq. : average number of behaviors per task; Cost : average share of per-task monetary cost attributable to the behavior.
Figure 2: Temporal distribution of the cost-inefficient behaviors. Bold curves aggregate configurations; thin curves show individual ones.
SubRetrv Actions
Occurrence Scenario
CC
MSA S46
MSA MM3
MSA Q35+
Cross-Agent SubRetrv
324 50.15%
—
—
—
Patch-Adjacent SubRetrv
56 8.67%
274 26.47%
277 17.96%
374 19.56%
Within-Sequence SubRetrv
171 26.47%
311 30.05%
255 16.54%
897 46.91%
Long-Distance SubRetrv
73 11.30%
275 26.57%
808 52.40%
472 24.69%
Other
22 3.41%
175 16.91%
202 13.10%
169 8.84%
Table 2: SubRetrv actions by scenario, assigned in top-to-bottom priority order. Subscripts indicate the percentage of SubRetrv actions.
Figure 5
Figure 6: Changes in average task cost versus behavior-attributed cost, both relative to baseline task cost. The shaded quadrant indicates joint improvement, and the lines show the fitted linear trends.
Appendix figures & tables16 assets
Supplementary material from the paper’s appendix.
Appendix
Category
Label
Description
Retrieval
Retrieve \ _Structure
Searches or explores the repository structure or current location.
Retrieve \ _Code
Reads or searches source code files, except test files or documents.
Retrieve \ _Tests
Reads or searches test files specifically.
Retrieve \ _Docs
Reads documentation (non-code references such as README).
Modification
Generate \ _Patch
Modifies a file inside the working tree (e.g., / testbed ) to fix the issue.
Generate \ _Tests
Writes tests or reproduction scripts (e.g., python - c ).
Appendix
Table 5: Action-label taxonomy.
Figure 7: Distribution of action labels over the RQ1 trajectories.
SWE-bench Verified (500 tasks)
SWE-bench Pro (731 public tasks)
Repository
Analysis
Eval.
Repository
Pool
Eval.
astropy/astropy
13
9
ansible/ansible
96
12
django/django
139
92
element-hq/element-web
56
8
matplotlib/matplotlib
20
14
flipt-io/flipt
85
10
mwaskom/seaborn
1
1
future-architect/vuls
62
9
pallets/flask
1
0
gravitational/teleport
76
10
Appendix
Table 6: Task selection by repository on SWE-bench Verified and SWE-bench Pro.
Pass@1
Cost
CoP
Min. robust ∣Δ∣ , R=1
Benchmark
Config
Run 1 / 2 / 3 (%)
spass (pp)
Run 1 / 2 / 3 ($)
scost (%)
sCoP (%)
Cost
Pass@1 (pp)
Verified-200
CC
74.00 / 75.50 / 75.50
0.87
0.598 / 0.525 / 0.515
8.35
9.56
15.87%
1.65
Verified-200
MSA S46
72.00 / 71.50 / 74.00
1.32
0.674 / 0.605 / 0.636
5.40
5.49
10.25%
2.51
Verified-200
MSA MM3
78.00 / 71.50 / 76.00
3.33
0.463 / 0.442 / 0.477
3.80
2.85
7.21%
6.33
Verified-200
MSA Q35+
65.00 / 70.50 / 66.00
2.93
0.123 / 0.110 / 0.115
5.70
9.68
10.82%
5.57
Pro-100
CC
55.00 / 56.00 / 52.00
2.08
0.917 / 0.922 / 0.881
2.44
1.94
4.64%
3.95
Appendix
Table 7: Noise floor of the eight baselines. The floor s is the standard deviation of Pass@1 (pp) and the coefficient of variation of cost and CoP (%). The last two columns give the smallest single-run ∣Δ∣ classified as robust ( 1.90s ). Per-run cost is averaged over the tasks completed in all three runs.
Pass@1 (pp)
Cost ($)
CoP
Benchmark
Config
Approach
Per run
sappr / sbase
Ratio
Per run
sappr / sbase
Ratio
Ratio
Verified-200
CC
DevSkills
73.00 / 73.50 / 74.00
0.50 / 0.87
0.58
0.497 / 0.487 / 0.435
7.06 / 8.35
0.84
0.80
Verified-200
CC
CodeGraph
73.50 / 73.50 / 75.50
1.15 / 0.87
1.33
0.590 / 0.613 / 0.587
2.36 / 8.35
0.28
0.36
Verified-200
MSA MM3
DevSkills
72.00 / 72.50 / 74.50
1.32 / 3.33
0.40
0.371 / 0.381 / 0.410
5.23 / 3.80
1.38
1.19
Verified-200
MSA Q35+
DevSkills
67.50 / 72.00 / 70.00
2.25 / 2.93
0.77
0.107 / 0.099 / 0.103
3.83 / 5.70
0.67
0.73
Verified-200
MSA Q35+
SynSkills
70.00 / 66.00 / 73.00
3.51 / 2.93
1.20
0.119 / 0.112 / 0.115
3.01 / 5.70
0.53
0.48
Appendix
Table 8: Approach cells executed three times to verify the equal-variance assumption. Each metric lists the approach and baseline noise ( sappr / sbase ) and their ratio. The remaining 15 approach cells are single runs.
Config
Pass@1
Cost ($)
CoP
CC
76.67%
0.520
0.679
MSA S46
74.33%
0.554
0.745
MSA MM3
72.67%
0.365
0.502
MSA Q35+
65.67%
0.102
0.156
Appendix
Table 9: Pass@1, cost, and CoP of the four agent configurations in RQ1 .
Metric
CC
MSA S46
MSA MM3
MSA Q35+
Avg. total lines read
1,032
573
912
1,051
Avg. lines re-read by SubRetrv
114
96
137
190
Overlapping rate
11.05%
16.69%
15.00%
18.10%
Appendix
Table 10: Total retrieved context and the lines re-read by SubRetrv actions.
Figure 8: Patch-Adjacent SubRetrv of MSA Q35+ on SWE-bench Verified task django -13158 .
Figure 9: Within-Sequence SubRetrv of MSA Q35+ on SWE-bench Verified task astropy -8707 .
Figure 10: Long-Distance SubRetrv of MSA MM3 on SWE-bench Verified task matplotlib -24627 .
Figure 11: File-based SimScrpt of MSA MM3 on SWE-bench Verified task matplotlib -24627 .
Figure 12: Distribution of ReTest actions per task.
Figure 14: Test-signal recovery of MSA MM3 on SWE-bench Verified task sympy -13878 .
Verified-200
Pro-100
CC
MSA S46
MSA MM3
MSA Q35+
CC
MSA S46
MSA MM3
MSA Q35+
CodeGraph task coverage (%)
91.50
95.00
97.50
93.17
77.00
85.00
77.00
74.00
Share of CodeGraph queries (%)
Read source
98.88
73.97
69.03
58.82
81.47
63.80
64.82
54.16
Locate symbol
0.45
20.65
18.07
36.48
7.99
17.04
19.46
31.51
Read callers
0.56
3.62
4.28
3.30
1.60
3.24
5.36
6.29
Appendix
Table 12: CodeGraph usage in RQ2 covering all 100 Pro tasks, including the percentage of tasks in which the agent issued at least one CodeGraph query, and the share of CodeGraph queries by functionality.
Tool
Returns
get_skeleton
Step indices, action labels, and truncated actions for a trajectory overview.
get_step([ids])
Full thoughts, commands, and command outputs for selected steps.
get_gt_patch
The ground-truth patch of the task.
get_agent_submission
The agent’s final submitted patch.
Appendix
Table 13: Toolset of the trajectory-analysis agent.
Config
Agent ($)
One-time ($)
Merge ($)
Total ($)
Task cost ($)
Ratio
CC
20.33
1.78
0.65
22.76
156.00
14.59%
MSA S46
34.48
2.03
1.08
37.59
166.06
22.64%
MSA MM3
3.72
1.65
0.49
5.86
109.50
5.35%
MSA Q35+
2.92
1.23
0.12
4.27
30.60
13.95%
Appendix
Table 14: Cost of synthesizing skills. Agent , One-time , and Merge are the total costs of the agent analysis, the single-call analysis, and the hierarchical consolidation; Total : the total cost of synthesizing skills; Task cost : the cost of solving all 300 RQ1 tasks; Ratio : Total / Task cost.
Modern coding agents expose multiple tool surfaces -- IDE primitives, bash, and Model Context Protocol (MCP) code-execution -- and the field has shipped three contradictory claims about which one matters. We run the missing crossed comparison: an integrity-clean three-arm ablation (baseline / bash_only / code_only) on synthetic computation tasks and SWE-bench Mini modification tasks, holding model, harness, and prompts fixed, with two agents (Claude Code, OpenAI Codex CLI) so the comparison spans both regime and agent-design axes. Across the four resulting (regime, agent) cells, restricting the agent to a single execute_code MCP tool is cheaper than -- or statistically tied with -- its cheapest tool-rich rival in three cells (significantly on Artifact/Claude and SWE-bench/Codex; directionally on Artifact/Codex), with pass rates statistically tied within each cell. The lone exception is SWE-bench/Claude, where code_only is directionally costlier (+14.4%, not significant); a conditional-cost analysis localizes that gap to failure-cost on doomed-run trajectories, not a per-edit tax on successful runs. Two implications: the cheapest tool surface is jointly determined by task regime and agent design rather than by either axis alone, and the headline cost signal lives in cache-adjusted cost -- not pass rate, which is invariant across surfaces at the model sizes we evaluate. The benchmark harness, task suite, and analysis code are available at https://github.com/hyang0129/onlycodes.
Hong Yang, Qi Yu, Travis Desell
Rochester Institute of Technology Rochester, New York, USA
Two prompts can request the same code change and produce the same correct patch, yet cause a coding agent to perform radically different kinds and amounts of work. We study this effect in a preregistered benchmark spanning 4,644 valid runs, 24 deterministic coding tasks, seven reasoning models, and two real agent harnesses. The central finding is that prompt wording does not merely scale total effort; it changes where that effort is spent. Multiple approaches and deep thinking primarily inflate reasoning. Multiple approaches increases reasoning by 2.4x to 7.4x across all six open models and creates about three elaborated but discarded solution branches, while still yielding only one implemented solution and no success gain. Maximum certainty activates a different pathway: repeated verification propagates into extra test runs, tool calls, turns, latency, and context growth. Runs with high redundant verification cost 18x the clean-run median, execute 2.5x more tool calls, and take 3x longer, again without a success gradient. These mechanisms therefore have distinct cost carriers: some prompts are reasoning-heavy and token-borne, while others are tool-heavy and system-borne. Harness design amplifies both effects and changes cost per successful task by 5x to 30x in our setting. The findings survive a frozen holdout, paraphrase tests, a Kimi-K3 replication, and a first-party Claude Sonnet 5 study. In contrast, bounded-efficiency wording preserves diagnosis and final validation while avoiding the measured waste mechanisms. Prompt engineering for coding agents is therefore work design: it determines what the agent thinks through, what it executes, and when it stops.
Coding agents increasingly operate in executable environments where a failed attempt produces actionable feedback rather than merely an incorrect answer. Existing cost-aware systems typically treat such failures as cascade decisions: try a cheap model first, then escalate hard cases to a stronger and more expensive model. In coding, however, execution feedback can also make further cheap-model recovery worthwhile, raising a budgeted deployment question: when should an agent spend more cheap compute, and when should it escalate? We formulate this post-failure decision as recovery routing over heterogeneous actions and train a supervised router from execution rollouts. To make the same router usable under changing budgets, we add a Conformal Risk Control (CRC) layer that selects a deployment-time cost penalty without retraining and provides marginal expected-cost control under exchangeability. Across held-out failures from five coding benchmarks, cheap recovery and escalation exhibit complementary success patterns. The calibrated frontier improves over fixed actions, prompt-only routers, and a binary cascade baseline; in the main GPT-5.4-nano/GPT-5.4 setting, one CRC-calibrated frontier point exceeds always-escalate solve rate while using 35% of its mean recovery cost. Code is available at https://github.com/Qijia-He/agent-budget-control.
Qijia He, Jiayi Cheng, Chenqian Le +8
University of Washington · New York University · ByteDance +1