Analyzing and Mitigating Cost-Inefficient Behaviors in Coding Agents
Organizations: Purdue University
Abstract
Although effective, coding agents often incur substantial monetary costs. Their recurring cost-inefficient behaviors remain underexplored. We conduct the first study of behavioral cost inefficiencies in coding agents, analyzing 1,200 trajectories from Claude Code and Mini-SWE-Agent across four configurations on SWE-bench Verified. We identify three cost-inefficient behaviors: subsumed retrieval, similar script generation, and test re-execution. We then evaluate three mitigation strategies: structure-aware retrieval, agent-synthesized skills, and developer-designed skills, over 10k trajectories on held-out SWE-bench Verified and Pro tasks. Our main findings are: (1) The three behaviors affect 79.00%--98.00% of coding tasks and account for up to 22.75% of task cost. (2) Structure-aware retrieval can introduce retrieval overhead and alter agent delegation, causing inconsistent improvements in retrieval efficiency and cost increases of up to 28.14%. (3) Agent-synthesized skills tend to produce low-level, trace-specific guidance, limiting their effectiveness and generality. (4) In contrast, developer-designed skills provide high-level, trace-agnostic guidance, reducing cost by up to 41.73%, roughly twice the maximum gain from agent-synthesized skills.
Figures & tables
| CC | MSA S46 | MSA MM3 | MSA Q35+ | |||||||||
| Behaviors | Tasks | Freq. | Cost | Tasks | Freq. | Cost | Tasks | Freq. | Cost | Tasks | Freq. | Cost |
| SubRetrv | 64.33% | 2.15 | 5.01% | 87.33% | 3.45 | 8.42% | 92.33% | 5.14 | 7.88% | 89.33% | 6.37 | 11.41% |
| SimScrpt | 20.67% | 0.43 | 1.02% | 51.33% | 2.54 | 9.57% | 68.00% | 4.29 | 8.85% | 57.67% | 3.20 | 7.91% |
| ReTest | 49.67% | 2.09 | 0.83% | 69.00% | 2.51 | 3.17% | 83.00% | 5.29 | 5.39% | 66.00% | 2.10 | 3.43% |
| Total | 79.00% | 4.67 | 6.86% | 97.33% | 8.50 | 21.16% | 98.00% | 14.72 | 22.12% | 96.67% | 11.68 | 22.75% |
| SubRetrv Actions | ||||
| Occurrence Scenario | CC | MSA S46 | MSA MM3 | MSA Q35+ |
| Cross-Agent SubRetrv | 324 50.15% | — | — | — |
| Patch-Adjacent SubRetrv | 56 8.67% | 274 26.47% | 277 17.96% | 374 19.56% |
| Within-Sequence SubRetrv | 171 26.47% | 311 30.05% | 255 16.54% | 897 46.91% |
| Long-Distance SubRetrv | 73 11.30% | 275 26.57% | 808 52.40% | 472 24.69% |
| Other | 22 3.41% | 175 16.91% | 202 13.10% | 169 8.84% |
Appendix figures & tables16 assets
Supplementary material from the paper’s appendix.
Appendix
| Category | Label | Description |
| Retrieval | Retrieve \ _Structure | Searches or explores the repository structure or current location. |
| Retrieve \ _Code | Reads or searches source code files, except test files or documents. | |
| Retrieve \ _Tests | Reads or searches test files specifically. | |
| Retrieve \ _Docs | Reads documentation (non-code references such as README). | |
| Modification | Generate \ _Patch | Modifies a file inside the working tree (e.g., / testbed ) to fix the issue. |
| Generate \ _Tests | Writes tests or reproduction scripts (e.g., python - c ). |
| SWE-bench Verified (500 tasks) | SWE-bench Pro (731 public tasks) | ||||
| Repository | Analysis | Eval. | Repository | Pool | Eval. |
| astropy/astropy | 13 | 9 | ansible/ansible | 96 | 12 |
| django/django | 139 | 92 | element-hq/element-web | 56 | 8 |
| matplotlib/matplotlib | 20 | 14 | flipt-io/flipt | 85 | 10 |
| mwaskom/seaborn | 1 | 1 | future-architect/vuls | 62 | 9 |
| pallets/flask | 1 | 0 | gravitational/teleport | 76 | 10 |
| Pass@1 | Cost | CoP | Min. robust , | |||||
| Benchmark | Config | Run 1 / 2 / 3 (%) | (pp) | Run 1 / 2 / 3 ($) | (%) | (%) | Cost | Pass@1 (pp) |
| Verified-200 | CC | 74.00 / 75.50 / 75.50 | 0.87 | 0.598 / 0.525 / 0.515 | 8.35 | 9.56 | 15.87% | 1.65 |
| Verified-200 | MSA S46 | 72.00 / 71.50 / 74.00 | 1.32 | 0.674 / 0.605 / 0.636 | 5.40 | 5.49 | 10.25% | 2.51 |
| Verified-200 | MSA MM3 | 78.00 / 71.50 / 76.00 | 3.33 | 0.463 / 0.442 / 0.477 | 3.80 | 2.85 | 7.21% | 6.33 |
| Verified-200 | MSA Q35+ | 65.00 / 70.50 / 66.00 | 2.93 | 0.123 / 0.110 / 0.115 | 5.70 | 9.68 | 10.82% | 5.57 |
| Pro-100 | CC | 55.00 / 56.00 / 52.00 | 2.08 | 0.917 / 0.922 / 0.881 | 2.44 | 1.94 | 4.64% | 3.95 |
| Pass@1 (pp) | Cost ($) | CoP | |||||||
| Benchmark | Config | Approach | Per run | / | Ratio | Per run | / | Ratio | Ratio |
| Verified-200 | CC | DevSkills | 73.00 / 73.50 / 74.00 | 0.50 / 0.87 | 0.58 | 0.497 / 0.487 / 0.435 | 7.06 / 8.35 | 0.84 | 0.80 |
| Verified-200 | CC | CodeGraph | 73.50 / 73.50 / 75.50 | 1.15 / 0.87 | 1.33 | 0.590 / 0.613 / 0.587 | 2.36 / 8.35 | 0.28 | 0.36 |
| Verified-200 | MSA MM3 | DevSkills | 72.00 / 72.50 / 74.50 | 1.32 / 3.33 | 0.40 | 0.371 / 0.381 / 0.410 | 5.23 / 3.80 | 1.38 | 1.19 |
| Verified-200 | MSA Q35+ | DevSkills | 67.50 / 72.00 / 70.00 | 2.25 / 2.93 | 0.77 | 0.107 / 0.099 / 0.103 | 3.83 / 5.70 | 0.67 | 0.73 |
| Verified-200 | MSA Q35+ | SynSkills | 70.00 / 66.00 / 73.00 | 3.51 / 2.93 | 1.20 | 0.119 / 0.112 / 0.115 | 3.01 / 5.70 | 0.53 | 0.48 |
| Config | Pass@1 | Cost ($) | CoP |
| CC | 76.67% | 0.520 | 0.679 |
| MSA S46 | 74.33% | 0.554 | 0.745 |
| MSA MM3 | 72.67% | 0.365 | 0.502 |
| MSA Q35+ | 65.67% | 0.102 | 0.156 |
| Metric | CC | MSA S46 | MSA MM3 | MSA Q35+ |
| Avg. total lines read | 1,032 | 573 | 912 | 1,051 |
| Avg. lines re-read by SubRetrv | 114 | 96 | 137 | 190 |
| Overlapping rate | 11.05% | 16.69% | 15.00% | 18.10% |
| Verified-200 | Pro-100 | |||||||
| CC | MSA S46 | MSA MM3 | MSA Q35+ | CC | MSA S46 | MSA MM3 | MSA Q35+ | |
| CodeGraph task coverage (%) | 91.50 | 95.00 | 97.50 | 93.17 | 77.00 | 85.00 | 77.00 | 74.00 |
| Share of CodeGraph queries (%) | ||||||||
| Read source | 98.88 | 73.97 | 69.03 | 58.82 | 81.47 | 63.80 | 64.82 | 54.16 |
| Locate symbol | 0.45 | 20.65 | 18.07 | 36.48 | 7.99 | 17.04 | 19.46 | 31.51 |
| Read callers | 0.56 | 3.62 | 4.28 | 3.30 | 1.60 | 3.24 | 5.36 | 6.29 |
| Tool | Returns |
| get_skeleton | Step indices, action labels, and truncated actions for a trajectory overview. |
| get_step([ids]) | Full thoughts, commands, and command outputs for selected steps. |
| get_gt_patch | The ground-truth patch of the task. |
| get_agent_submission | The agent’s final submitted patch. |
| Config | Agent ($) | One-time ($) | Merge ($) | Total ($) | Task cost ($) | Ratio |
| CC | 20.33 | 1.78 | 0.65 | 22.76 | 156.00 | 14.59% |
| MSA S46 | 34.48 | 2.03 | 1.08 | 37.59 | 166.06 | 22.64% |
| MSA MM3 | 3.72 | 1.65 | 0.49 | 5.86 | 109.50 | 5.35% |
| MSA Q35+ | 2.92 | 1.23 | 0.12 | 4.27 | 30.60 | 13.95% |