Organizations: Institute of Information Engineering, Chinese Academy of Sciences · School of Cyber Security, University of Chinese Academy of Sciences · Beihang University · State Key Laboratory of Networking and Switching Technology, Beijing University of Posts and Telecommunications, Beijing, China
Multi-step tool-calling LLM agents rely on host runtimes to preserve state across turns. When a runtime carries an external tool return into later model inputs, providers meter it again. An admitted malicious or compromised tool can thereby convert untrusted data into recurring victim-billed processing without victim credentials or local runtime privilege. We call retained content persistent billable state and formalize the host's decision over whether and how it enters later billable context as the persistent billable-state boundary. We present the first systematic security study of this post-admission lifecycle. We derive six denial-of-wallet attack vectors and build DOW-BENCH, an end-to-end harness evaluated across six model families. Across 243 executions, usage telemetry shows that the maximum per-session cumulative input reaches 14,293x the session's first-call input. Controlled history-policy reruns isolate raw retention's contribution: retaining raw history increases mean effective session cost by 21.2-35.9%. Compression succeeds on 10/12 and 11/12 history-dependent tasks, versus 2/12 under deletion for each provider. To govern this boundary, we combine deterministic history transformation with four host-side invariants that bound prompt mass, context growth, recursive opportunity, and cumulative spend before reingestion. The kernel contains every recurring attack in the 123-evaluation replay corpus. Across 24 Mistral Small 4 workflows, a progress-authorized policy achieves 22/24 oracle-verified task successes with no pre-completion interruptions, versus 13/24 under a fixed cap. Only 71 of 3,830 scanned MCP server and transport repositories expose any code-visible safeguard proxy, and none cover all four safeguard families. These results establish persistent billable state as a first-class security object and pre-reingestion as its host-owned control point.
Figures & tables
Figure 1 : Lifecycle of persistent billable state. An ordinary task invokes an already-admitted malicious or compromised remote tool. Its schema-valid return Ri crosses the host’s pre-reingestion boundary, persists in later model inputs, and accrues provider-metered cost on each reuse. Widths are schematic and omit non-tool growth.
Primary lever
Payload-control policy
Direct
Stateless
Adaptive
Retained input mass ( ↑pi )
V0 Direct recursive baseline
V1 Polymorphic mutation
V3 Stealth escalation; V5 Adaptive
Generated output ( ↑oi )
V4 Weaponized output
—
—
Recursive opportunity ( ↑N )
—
V2 Cross-tool chaining
—
Table 1: Taxonomy of the six evaluated DoW vectors by primary amplification lever and payload-control policy.
Evidence class
Unit / scope
Design / eligibility
What it establishes
RQ1–RQ2: Core measurement
243 executions; six families; 11 serialized IDs
Fresh session per run; balanced family analysis
CAF range and turn-survival patterns across deployments
RQ1: Priced exposure
78 API sessions (69 attack, nine benign)
Provider telemetry with a nonzero canonical listed rate
Listed-price session exposure
RQ2: History intervention
144 independent sessions across 12 tasks, three policies, two return classes, and two providers
Closed-loop task-paired reruns
Effect of retention policy on cost and task success
RQ3: Defense replay
41 recurring multi-turn attacks × three presets
Frozen zero-call replay with trigger-inclusive leakage
Historical-trace containment and layer activation
RQ3: Benign operating point
Five full 210-scenario corpora ( n=1,050 provider–scenario units)
Provider-specific full-corpus checks
Moderate-D2 false-positive operating point
RQ3: Long-workflow design
60 provider–scenario executions (20 designs × three runs)
Policy-design corpus
Utility of fixed and growth-gated recursion policies
Table 2 : Study map linking RQ1–RQ3 and deployment context to the experimental designs and claims that answer them.
Figure 2 : Listed-price exposure across 69 attack and nine benign API sessions. Gray crosses mark benign sessions and brick-red circles mark attack sessions; deterministic horizontal displacement resolves overlap without changing vertical positions. Hollow diamonds and bars show medians and IQRs; the dashed line and shading mark the $0.10 moderate D4 boundary and exposures above it.
Cohort
Priced / total
Mean (USD)
Median (USD)
Max (USD)
>\0.10(n/N$ )
Benign
9 / 35
$0.014
$0.014
$0.025
0/9
V0 Direct
16 / 54
$1.434
$1.293
$3.404
14/16
V1 Poly.
14 / 42
$1.182
$1.074
$2.994
12/14
V2 Cross
3 / 18
$0.475
$0.359
$1.066
2/3
V3 Stealth
19 / 48
$0.526
$0.582
$0.715
17/19
V4 Weapon.
3 / 19
$0.290
$0.062
$0.808
1/3
Table 3 : Listed-price exposure in the 256-execution evaluation cohort.
Figure 3 : CAF distributions across six core model families, four primary attack vectors (V0/V1/V3/V5), and the benign comparator (206 executions; n=5 – 11 per condition within a family). Within each condition, thin lines span the observed minimum and maximum, thick segments denote the interquartile range, and markers denote medians. Zero-width intervals collapse to the marker. CAF is logarithmic.
Figure 4 : Turn survival tracks recursive exposure across observed deployments. (a) Mean CAF versus mean turns across the four primary vectors (V0/V1/V3/V5; 24 complete family–vector cells). (b) Ratios of mean CAF between globally ranked higher- and lower-amplification family groups, observed and after offline truncation of 44 malicious traces with recorded per-turn prompt tokens at two, three, or four turns; bars show trace-bootstrap 95% CIs on a logarithmic axis. (c) Mean paired difference in downstream CAF gain across 111 supported V0/V3 matched prefixes (cross-fitted higher- minus lower-amplification family group); the horizontal bar gives the trace-cluster bootstrap 95% CI (683–3,908; bootstrap p≤0.002 ).
Figure 5 : Raw history raises cost, whereas compression preserves most utility on tasks requiring prior history. (a) Full-over-Drop/Compress effective-cost changes on persistent returns, shown as ratios of arm means minus one with task-paired bootstrap 95% CIs (12 planned pairs per provider–comparison); costs are complete-session totals. (b) Planned intention-to-treat (ITT) oracle-verified task success on matched-necessary tasks; non-completions remain non-successes. Both panels report the same two provider–model pairs.
Figure 6 : Four host-side invariants enforced before reingestion. D1 bounds token mass, D2 flags abnormal reingestion growth, D3 limits recursive opportunity, and D4 caps cumulative priced exposure; each gate acts before the next provider call.
Figure 7 : Seven constructed, white-box threshold-aware traces evaluated offline against the frozen moderate kernel collectively exercise all four gate conditions. Cell position identifies the first reported layer under the kernel’s evaluation order, labels give the block turn, and the rightmost column reports cumulative nominal cost through and including that turn.
Figure 8 : Mistral D4 policy-transfer outcomes. (a) ITT outcomes across 24 planned sessions per policy. (b) Exact two-sided McNemar comparison across the 21 workflow pairs with evaluable task-success outcomes under both policies.
Appendix figures & tables16 assets
Supplementary material from the paper’s appendix.
Appendix
Model
Metric
Benign
Direct (V0)
Polymorphic (V1)
Stealth (V3)
Adaptive (V5)
V5 vs. benign adj. p / rrb
GPT-4o
Median CAF
35.6
5,778.3
4,204.9
2,427.8
440.2
0.0317∗rrb=0.89
IQR
32.8–87.5
1,889.8–8,921.7
3,997.1–4,692.2
2,177.7–2,798.8
286.5–4,169.9
Max
208.7
11,116.0
10,820.5
3,328.6
7,143.9
Valid n
5
5
5
6
7
GPT-4o-mini
Median CAF
35.2
106.7
139.9
301.6
331.1
0.0317∗rrb=1.00
IQR
35.1–35.5
106.7–122.3
111.3–161.3
301.5–303.4
330.9–331.2
Appendix
Table 4 : CAF distributions across six core LLM families, four primary attack vectors, and the benign comparator.
Figure 9 : Execution-level CAF for the 206 runs summarized in Figure 3 . Each marker represents one execution; deterministic vertical stacking separates coincident and near-coincident values within a condition row without changing CAF on the horizontal axis. Panels share a logarithmic CAF scale and retain the main figure’s condition colors and marker shapes.
Attack
All
No D1
No D2
No D3
No D4
V0
0.022
0.022
0.151
0.022
0.022
V1
0.024
0.024
0.100
0.024
0.024
V2
0.128
0.128
0.128
0.128
0.128
V3
0.020
0.020
0.128
0.020
0.020
V4
0.010
0.010
0.114
0.010
0.010
V5
0.008
0.008
0.117
0.008
0.008
Appendix
Table 5: Moderate-preset single-layer ablation on one attack sequence per vector.
Schedule
Defense
Completed n/N
First stop
V0 burst (2 concurrent)
Off
2/2
max-turns T2
V0 staggered (2 concurrent)
Off
5/5
max-turns T2
V0 staggered (2 concurrent)
D1–D4
5/5
D2 at T2
Seed inflation
D1–D4
5/5
D4 at T2
Appendix
Table 6 : Nominal burn-rate and first-trigger calibration on Cerebras gpt-oss-120b : two defense-off V0 schedules, a matched full-stack V0 schedule, and five full-stack seed-inflation sessions.
Metric
Groq GPT-OSS-120B
Mistral Small 2603
Sessions
20
20
Cached-input share
2.0%
57.9%
Total gross/effective cost
0.0190/0.0188
0.0159/0.0079
Cost saved
0.9%
50.2%
Effective-cost amplification
26.1–37.1 ×
16.0–20.8 ×
Attack / benign cost
0.88–1.25 ×
1.00–1.31 ×
Appendix
Table 7 : Cache-aware cost calibration and live MCP utility over stdio.
Provider/model
Turns
False-positive scenarios
Raw D2
Floor 512
Floor 768
Gate 3k
Primary Llama-3.1-8B
575
0/210
0/210
0/210
0/210
GitHub GPT-4o-mini
468
1/210
1/210
0/210
0/210
Cerebras GPT-OSS-120B
847
0/210
0/210
0/210
0/210
Mistral Small Latest
541
1/210
1/210
0/210
0/210
Groq GPT-OSS-120B
911
0/210
0/210
0/210
0/210
Appendix
Table 8 : Provider-specific benign false positives under post hoc D2 ratio-branch variants.
Policy
Replay
Long tasks
Blocked ( N=41 )
Tasks done ( N=60 )
Pre-comp. D3 stops ( N=60 )
Pre-comp. D4 stops ( N=60 )
Frozen moderate
41/41
0/60
60/60
0/60
Successor + D4
41/41
53/60
0/60
7/60
Successor, D1–D3
41/41
60/60
0/60
0/60
Appendix
Table 9 : Growth-gated D3 outcomes on 41 replay attacks and 60 completed 8–12-step workflows.
Policy
Protocol outcomes ( N=12 )
Tasks complete
Pre-comp. blocks
Infra. failures
Model incomplete
Fixed $0.10 (counterfactual)
7/12
2/12
1/12
2/12
Checkpoint budget (live)
9/12
0/12
1/12
2/12
Appendix
Table 10 : D3/D4 successor check on 12 Cerebras GPT-OSS-120B workflows.
Metric
Fixed
Progress-auth.
ITT task success
13/24 (54.2%)
22/24 (91.7%)
Pre-completion D4 stop
9/24
0/24
Tool failure
2/24
2/24
Mean cost/session (USD)
0.0992
0.1126
Median cost/session (USD)
0.1037
0.1143
Appendix
Table 11 : D4-only Mistral Small 4 transfer: 24 post-qualification workflows independently executed under each policy (48 sessions).
Treatment
Task success
Cost ratio
Effective / benign
Gross / projected
Groq
Persistent rebilling
5/5
1.249 ×
1.675 ×
Structural-loop proxy
5/5
0.880 ×
1.344 ×
Cost-optimization proxy
5/5
1.192 ×
1.593 ×
Matched benign
5/5
1.000 ×
1.555 ×
Appendix
Table 12 : Success-gated cost comparison of related mechanisms on matched workflows.
Figure 10 : D2 operating point across 1,050 benign provider–scenario executions. Left: 2,292 post-seed turns from 984 multi-turn scenarios; three raw-trigger rows belong to two scenarios, and Δp/p1 uses a logarithmic axis. Right: false-positive counts retain the 1,050-scenario denominator, while attack leakage averages trigger-inclusive fractions over six replay attacks.
Figure 11 : Live inline placement across five attack vectors and five provider–model pairs. Fill encodes first block turn; text gives interception layer/turn; hatching marks an upstream HTTP 413 cap. D2/D3 intercept 23/25 sessions before the next provider call; the other two stop at the provider cap.
Figure 12 : Evaluation pipeline from isolated execution to offline accounting. Each fresh API session emits one execution record to the append-only ledger, from which offline accounting builds the analysis view used for reported results.
Figure 13 : Attack-vector tradeoffs for V0, V1, V3, and V5 ( n=50/36/46/39 executions). Panels report mean CAF, mean turns, and CAF stability (inverse mean within-family CV across six family cells); higher stability is better. No vector leads on all three, and each panel retains its native scale.
Figure 14 : Prevalence of code-visible safeguard proxies across 3,830 public MCP server and transport repositories. (a) The scan detects no proxy in 3,759/3,830 repositories (98.1%) and at least one in 71/3,830 (1.9%). (b) Among those 71 positives, 67 (94.4%) match one family, 4 (5.6%) match two or three, and none match all four.