Control the Harness, Control the Cost: Routing and Governing AI Coding Agents in the Enterprise
Organizations: Accenture Responsible AI · Accenture Americas Advanced AI Practice · Harvard Extension School, Harvard University
Abstract
Harnesses, the products that run AI coding agents, are multiplying, and enterprises are rolling them out to their employees: what started as pilots with a few hundred seats is scaling to tens of thousands. Most enterprises do not build these harnesses but buy them from large vendors, such as Anthropic's Claude Code or OpenAI's Codex. A harness decides which model answers, what the model reads, how the prompt cache is used and which subagents run, so it picks the rate on the price sheet and sets the volume bought at it. Enterprises that keep a proprietary or untuned harness at its defaults inherit these choices and their bill. We build a fast, customisable router in which Jev, a classifier with calibrated probabilities, labels every prompt against a bring-your-own taxonomy of agentic requests. Because one user turn is many requests over a prompt cache that belongs to one model, the router moves work only where no running conversation has to rebuild its cache: at session start, in side lanes and at subagent launch. From the price sheet we derive when a mid-task switch pays back, and a crossover: on long tool-heavy sessions the highest-priced model costs less than the next tier, as repricing about 10,000 real sessions from public datasets confirms. In an emulated enterprise of 10,000 seats with user behaviour taken from these datasets, the router recovers 14 to 21% of model spend at Anthropic's list prices of 21 September 2026, $3.3M to $5.0M a year. The paper also maps the risks across twenty harnesses, prices the dependence on one vendor's models, and proposes a control plane that enterprises can run from within, starting now, with a ladder for deciding later whether to own the harness.
Figures & tables
| Vendor | Cache read | Cache write | Lifetime | Minimum prefix |
|---|---|---|---|---|
| Anthropic [ 11 , 10 ] | 0.025 (Fable 5.1), 0.1 others | 1.25 (5 min), 2 (1 h) | 5 min or 1 h | 512 to 4,096 tokens |
| OpenAI, GPT-5.6 and later [ 12 , 13 ] | 0.1 | 1.25 | 30 min | 1,024 tokens |
| Google, Gemini [ 14 ] | 0.1 | input price plus a storage fee per hour (explicit cache) | not stated (implicit caching on by default) | 2,048 to 4,096 tokens |
| Tool steps per user turn | Context above which Fable 5.1 is cheaper than Opus 5 |
|---|---|
| 3 | 385k |
| 6 | 211k |
| 12 | 105k |
| 18 | 65k |
| 30 | 32k |
| Quantity | Value |
|---|---|
| Scenarios / turns / job types | 46 / 109 / 11 |
| Classifier accuracy: action / domain / complexity / history | 74% / 80% / 63% / 82% |
| False downgrades among moved turns / held in place (of routable) | 12.5% / 42% |
| As-is spend, 10,000 seats | 23.7M a year |
| Routed, today’s labels | -$ 13.8%) |
| Routed, perfect labels | -$ 21.1%) |
| Case (population, same labels) | As-is / month | Routed / month | Saving |
|---|---|---|---|
| Base case | $1.97M | $1.70M | 13.8% |
| Ledger rebase at task boundaries | $1.97M | $1.70M | 13.9% |
| One-hour cache lifetime (as-is and routed) | $1.93M | $1.60M | 17.1% |
| No Fable 5.1 licence (as-is on Opus 5, router capped at Opus 5) | $2.01M | $1.97M | 2.3% |
| Everyone as-is on Opus 5 (Fable 5.1 licensed for the router) | $2.01M | $1.70M | 15.6% |
| Confidence threshold 0.7 (base 0.8) | $1.97M | $1.78M | 9.8% |
| Always-loaded prefix per request | Fable 5.1 | Opus 5 | Sonnet 5 |
|---|---|---|---|
| 15k: lean default (system prompt and core tools) | $0.90 | $1.51 | $0.60 |
| 25k: default plus a 10k instruction and skills bundle | $1.49 | $2.52 | $1.01 |
| 55k: observed power-user configuration with plugins | $3.29 | $5.54 | $2.22 |
| 159k: the same if deferred tool schemas were loaded eagerly | $9.50 | $16.02 | $6.41 |
| 10,000 returns at 300k context | On the return | Each later request | Return vs re-send |
|---|---|---|---|
| Re-send the full history (default) | $18,750 | $1,500 | +0% |
| Prune to 30k first (clear old reasoning and tool results) | $1,875 | $150 | -90% |
| Compact to 30k, summary by Opus 5 | $20,625 | $150 | +10% |
| Compact to 30k, summary by Sonnet 5 | $9,375 | $150 | -50% |
| Rung | What it buys | What it costs |
|---|---|---|
| 0. Vendor defaults | Nothing to run | Every risk in Section 3 ; the highest bill |
| 1. Configured | The bundle above, model defaults, subagent default, permissions | Configuration per harness; no measurement |
| 2. Controlled | Gateway, session-start routing, side lanes via hooks, telemetry, policy floors, taxonomy; the 14 to 21% of Section 5.1 | A small platform team; classifier hosting; works-council and legal review of the policy |
| 3. Owned | An internal harness on a vendor SDK or an open-source base with pluggable models, native lanes, one bundle for everyone; other vendors’ models (Figure 10 ) | A product team; keeping pace with weekly vendor releases; the dependency moves to maintainers |