Harnesses, the products that run AI coding agents, are multiplying, and enterprises are rolling them out to their employees: what started as pilots with a few hundred seats is scaling to tens of thousands. Most enterprises do not build these harnesses but buy them from large vendors, such as Anthropic's Claude Code or OpenAI's Codex. A harness decides which model answers, what the model reads, how the prompt cache is used and which subagents run, so it picks the rate on the price sheet and sets the volume bought at it. Enterprises that keep a proprietary or untuned harness at its defaults inherit these choices and their bill. We build a fast, customisable router in which Jev, a classifier with calibrated probabilities, labels every prompt against a bring-your-own taxonomy of agentic requests. Because one user turn is many requests over a prompt cache that belongs to one model, the router moves work only where no running conversation has to rebuild its cache: at session start, in side lanes and at subagent launch. From the price sheet we derive when a mid-task switch pays back, and a crossover: on long tool-heavy sessions the highest-priced model costs less than the next tier, as repricing about 10,000 real sessions from public datasets confirms. In an emulated enterprise of 10,000 seats with user behaviour taken from these datasets, the router recovers 14 to 21% of model spend at Anthropic's list prices of 21 September 2026, $3.3M to $5.0M a year. The paper also maps the risks across twenty harnesses, prices the dependence on one vendor's models, and proposes a control plane that enterprises can run from within, starting now, with a ladder for deciding later whether to own the harness.
Figures & tables
Vendor
Cache read
Cache write
Lifetime
Minimum prefix
Anthropic [ 11 , 10 ]
0.025 × (Fable 5.1), 0.1 × others
1.25 × (5 min), 2 × (1 h)
5 min or 1 h
512 to 4,096 tokens
OpenAI, GPT-5.6 and later [ 12 , 13 ]
0.1 ×
1.25 ×
30 min
1,024 tokens
Google, Gemini [ 14 ]
0.1 ×
input price plus a storage fee per hour (explicit cache)
not stated (implicit caching on by default)
2,048 to 4,096 tokens
Table 1: Prompt-cache price rules by vendor (Anthropic’s of 21 September 2026, OpenAI’s and Google’s of 23 September 2026). Multipliers are of the model’s input price.
Tool steps per user turn
Context above which Fable 5.1 is cheaper than Opus 5
3
385k
6
211k
12
105k
18
65k
30
32k
Table 2: Context above which Fable 5.1 is cheaper than Opus 5 for one warm turn, by tool steps per turn (assumptions as in Figure 2 ).
Quantity
Value
Scenarios / turns / job types
46 / 109 / 11
Classifier accuracy: action / domain / complexity / history
74% / 80% / 63% / 82%
False downgrades among moved turns / held in place (of routable)
12.5% / 42%
As-is spend, 10,000 seats
1,974kamonth,23.7M a year
Routed, today’s labels
1,701kamonth(-$ 13.8%)
Routed, perfect labels
1,558kamonth(-$ 21.1%)
Table 3: Case-study summary (list prices of 21 September 2026; 10,000 seats; live classifier labels, cached).
Case (population, same labels)
As-is / month
Routed / month
Saving
Base case
$1.97M
$1.70M
13.8%
Ledger rebase at task boundaries
$1.97M
$1.70M
13.9%
One-hour cache lifetime (as-is and routed)
$1.93M
$1.60M
17.1%
No Fable 5.1 licence (as-is on Opus 5, router capped at Opus 5)
$2.01M
$1.97M
2.3%
Everyone as-is on Opus 5 (Fable 5.1 licensed for the router)
$2.01M
$1.70M
15.6%
Confidence threshold 0.7 (base 0.8)
$1.97M
$1.78M
9.8%
Table 4: Sensitivities (population, monthly, same labels).
Always-loaded prefix per request
Fable 5.1
Opus 5
Sonnet 5
15k: lean default (system prompt and core tools)
$0.90
$1.51
$0.60
25k: default plus a 10k instruction and skills bundle
$1.49
$2.52
$1.01
55k: observed power-user configuration with plugins
$3.29
$5.54
$2.22
159k: the same if deferred tool schemas were loaded eagerly
$9.50
$16.02
$6.41
Table 5: Cost per session of the always-loaded prefix (ten turns of eighteen tool steps, warm cache; list prices of 21 September 2026). Every token in the prefix is re-read on every request; the 55k row is a measured power-user configuration.
10,000 returns at 300k context
On the return
Each later request
Return vs re-send
Re-send the full history (default)
$18,750
$1,500
+0%
Prune to 30k first (clear old reasoning and tool results)
$1,875
$150
-90%
Compact to 30k, summary by Opus 5
$20,625
$150
+10%
Compact to 30k, summary by Sonnet 5
$9,375
$150
-50%
Table 6: Returning to a conversation whose cache has expired: 10,000 returns at 300k tokens of context on Opus 5 (list prices of 21 September 2026; 5-minute cache). Pruning clears old reasoning and tool results without a model call; a summary is written by the named model, which reads the full history uncached. The 30k is 15k of system prompt and tools plus a 15k-token summary. “Each later request” is the cache read of the history on every request after the return.
Rung
What it buys
What it costs
0. Vendor defaults
Nothing to run
Every risk in Section 3 ; the highest bill
1. Configured
The bundle above, model defaults, subagent default, permissions
Configuration per harness; no measurement
2. Controlled
Gateway, session-start routing, side lanes via hooks, telemetry, policy floors, taxonomy; the 14 to 21% of Section 5.1
A small platform team; classifier hosting; works-council and legal review of the policy
3. Owned
An internal harness on a vendor SDK or an open-source base with pluggable models, native lanes, one bundle for everyone; other vendors’ models (Figure 10 )
A product team; keeping pace with weekly vendor releases; the dependency moves to maintainers