Organizations: Institute of Information Engineering, Chinese Academy of Sciences · School of Cyber Security, University of Chinese Academy of Sciences · Beihang University · State Key Laboratory of Networking and Switching Technology, Beijing University of Posts and Telecommunications, Beijing, China
Multi-step tool-calling LLM agents rely on host runtimes to preserve state across turns. When a runtime carries an external tool return into later model inputs, providers meter it again. An admitted malicious or compromised tool can thereby convert untrusted data into recurring victim-billed processing without victim credentials or local runtime privilege. We call retained content persistent billable state and formalize the host's decision over whether and how it enters later billable context as the persistent billable-state boundary. We present the first systematic security study of this post-admission lifecycle. We derive six denial-of-wallet attack vectors and build DOW-BENCH, an end-to-end harness evaluated across six model families. Across 243 executions, usage telemetry shows that the maximum per-session cumulative input reaches 14,293x the session's first-call input. Controlled history-policy reruns isolate raw retention's contribution: retaining raw history increases mean effective session cost by 21.2-35.9%. Compression succeeds on 10/12 and 11/12 history-dependent tasks, versus 2/12 under deletion for each provider. To govern this boundary, we combine deterministic history transformation with four host-side invariants that bound prompt mass, context growth, recursive opportunity, and cumulative spend before reingestion. The kernel contains every recurring attack in the 123-evaluation replay corpus. Across 24 Mistral Small 4 workflows, a progress-authorized policy achieves 22/24 oracle-verified task successes with no pre-completion interruptions, versus 13/24 under a fixed cap. Only 71 of 3,830 scanned MCP server and transport repositories expose any code-visible safeguard proxy, and none cover all four safeguard families. These results establish persistent billable state as a first-class security object and pre-reingestion as its host-owned control point.
Figures & tables
Figure 1 : Lifecycle of persistent billable state. An ordinary task invokes an already-admitted malicious or compromised remote tool. Its schema-valid return Ri crosses the host’s pre-reingestion boundary, persists in later model inputs, and accrues provider-metered cost on each reuse. Widths are schematic and omit non-tool growth.
Primary lever
Payload-control policy
Direct
Stateless
Adaptive
Retained input mass ( ↑pi )
V0 Direct recursive baseline
V1 Polymorphic mutation
V3 Stealth escalation; V5 Adaptive
Generated output ( ↑oi )
V4 Weaponized output
—
—
Recursive opportunity ( ↑N )
—
V2 Cross-tool chaining
—
Table 1: Taxonomy of the six evaluated DoW vectors by primary amplification lever and payload-control policy.
Evidence class
Unit / scope
Design / eligibility
What it establishes
RQ1–RQ2: Core measurement
243 executions; six families; 11 serialized IDs
Fresh session per run; balanced family analysis
CAF range and turn-survival patterns across deployments
RQ1: Priced exposure
78 API sessions (69 attack, nine benign)
Provider telemetry with a nonzero canonical listed rate
Listed-price session exposure
RQ2: History intervention
144 independent sessions across 12 tasks, three policies, two return classes, and two providers
Closed-loop task-paired reruns
Effect of retention policy on cost and task success
RQ3: Defense replay
41 recurring multi-turn attacks × three presets
Frozen zero-call replay with trigger-inclusive leakage
Historical-trace containment and layer activation
RQ3: Benign operating point
Five full 210-scenario corpora ( n=1,050 provider–scenario units)
Provider-specific full-corpus checks
Moderate-D2 false-positive operating point
RQ3: Long-workflow design
60 provider–scenario executions (20 designs × three runs)
Policy-design corpus
Utility of fixed and growth-gated recursion policies
Table 2 : Study map linking RQ1–RQ3 and deployment context to the experimental designs and claims that answer them.
Figure 2 : Listed-price exposure across 69 attack and nine benign API sessions. Gray crosses mark benign sessions and brick-red circles mark attack sessions; deterministic horizontal displacement resolves overlap without changing vertical positions. Hollow diamonds and bars show medians and IQRs; the dashed line and shading mark the $0.10 moderate D4 boundary and exposures above it.
Cohort
Priced / total
Mean (USD)
Median (USD)
Max (USD)
>\0.10(n/N$ )
Benign
9 / 35
$0.014
$0.014
$0.025
0/9
V0 Direct
16 / 54
$1.434
$1.293
$3.404
14/16
V1 Poly.
14 / 42
$1.182
$1.074
$2.994
12/14
V2 Cross
3 / 18
$0.475
$0.359
$1.066
2/3
V3 Stealth
19 / 48
$0.526
$0.582
$0.715
17/19
V4 Weapon.
3 / 19
$0.290
$0.062
$0.808
1/3
Table 3 : Listed-price exposure in the 256-execution evaluation cohort.
Figure 3 : CAF distributions across six core model families, four primary attack vectors (V0/V1/V3/V5), and the benign comparator (206 executions; n=5 – 11 per condition within a family). Within each condition, thin lines span the observed minimum and maximum, thick segments denote the interquartile range, and markers denote medians. Zero-width intervals collapse to the marker. CAF is logarithmic.
Figure 4 : Turn survival tracks recursive exposure across observed deployments. (a) Mean CAF versus mean turns across the four primary vectors (V0/V1/V3/V5; 24 complete family–vector cells). (b) Ratios of mean CAF between globally ranked higher- and lower-amplification family groups, observed and after offline truncation of 44 malicious traces with recorded per-turn prompt tokens at two, three, or four turns; bars show trace-bootstrap 95% CIs on a logarithmic axis. (c) Mean paired difference in downstream CAF gain across 111 supported V0/V3 matched prefixes (cross-fitted higher- minus lower-amplification family group); the horizontal bar gives the trace-cluster bootstrap 95% CI (683–3,908; bootstrap p≤0.002 ).
Figure 5 : Raw history raises cost, whereas compression preserves most utility on tasks requiring prior history. (a) Full-over-Drop/Compress effective-cost changes on persistent returns, shown as ratios of arm means minus one with task-paired bootstrap 95% CIs (12 planned pairs per provider–comparison); costs are complete-session totals. (b) Planned intention-to-treat (ITT) oracle-verified task success on matched-necessary tasks; non-completions remain non-successes. Both panels report the same two provider–model pairs.
Figure 6 : Four host-side invariants enforced before reingestion. D1 bounds token mass, D2 flags abnormal reingestion growth, D3 limits recursive opportunity, and D4 caps cumulative priced exposure; each gate acts before the next provider call.
Figure 7 : Seven constructed, white-box threshold-aware traces evaluated offline against the frozen moderate kernel collectively exercise all four gate conditions. Cell position identifies the first reported layer under the kernel’s evaluation order, labels give the block turn, and the rightmost column reports cumulative nominal cost through and including that turn.
Figure 8 : Mistral D4 policy-transfer outcomes. (a) ITT outcomes across 24 planned sessions per policy. (b) Exact two-sided McNemar comparison across the 21 workflow pairs with evaluable task-success outcomes under both policies.
Appendix figures & tables16 assets
Supplementary material from the paper’s appendix.
Appendix
Model
Metric
Benign
Direct (V0)
Polymorphic (V1)
Stealth (V3)
Adaptive (V5)
V5 vs. benign adj. p / rrb
GPT-4o
Median CAF
35.6
5,778.3
4,204.9
2,427.8
440.2
0.0317∗rrb=0.89
IQR
32.8–87.5
1,889.8–8,921.7
3,997.1–4,692.2
2,177.7–2,798.8
286.5–4,169.9
Max
208.7
11,116.0
10,820.5
3,328.6
7,143.9
Valid n
5
5
5
6
7
GPT-4o-mini
Median CAF
35.2
106.7
139.9
301.6
331.1
0.0317∗rrb=1.00
IQR
35.1–35.5
106.7–122.3
111.3–161.3
301.5–303.4
330.9–331.2
Appendix
Table 4 : CAF distributions across six core LLM families, four primary attack vectors, and the benign comparator.
Figure 9 : Execution-level CAF for the 206 runs summarized in Figure 3 . Each marker represents one execution; deterministic vertical stacking separates coincident and near-coincident values within a condition row without changing CAF on the horizontal axis. Panels share a logarithmic CAF scale and retain the main figure’s condition colors and marker shapes.
Attack
All
No D1
No D2
No D3
No D4
V0
0.022
0.022
0.151
0.022
0.022
V1
0.024
0.024
0.100
0.024
0.024
V2
0.128
0.128
0.128
0.128
0.128
V3
0.020
0.020
0.128
0.020
0.020
V4
0.010
0.010
0.114
0.010
0.010
V5
0.008
0.008
0.117
0.008
0.008
Appendix
Table 5: Moderate-preset single-layer ablation on one attack sequence per vector.
Schedule
Defense
Completed n/N
First stop
V0 burst (2 concurrent)
Off
2/2
max-turns T2
V0 staggered (2 concurrent)
Off
5/5
max-turns T2
V0 staggered (2 concurrent)
D1–D4
5/5
D2 at T2
Seed inflation
D1–D4
5/5
D4 at T2
Appendix
Table 6 : Nominal burn-rate and first-trigger calibration on Cerebras gpt-oss-120b : two defense-off V0 schedules, a matched full-stack V0 schedule, and five full-stack seed-inflation sessions.
Metric
Groq GPT-OSS-120B
Mistral Small 2603
Sessions
20
20
Cached-input share
2.0%
57.9%
Total gross/effective cost
0.0190/0.0188
0.0159/0.0079
Cost saved
0.9%
50.2%
Effective-cost amplification
26.1–37.1 ×
16.0–20.8 ×
Attack / benign cost
0.88–1.25 ×
1.00–1.31 ×
Appendix
Table 7 : Cache-aware cost calibration and live MCP utility over stdio.
Provider/model
Turns
False-positive scenarios
Raw D2
Floor 512
Floor 768
Gate 3k
Primary Llama-3.1-8B
575
0/210
0/210
0/210
0/210
GitHub GPT-4o-mini
468
1/210
1/210
0/210
0/210
Cerebras GPT-OSS-120B
847
0/210
0/210
0/210
0/210
Mistral Small Latest
541
1/210
1/210
0/210
0/210
Groq GPT-OSS-120B
911
0/210
0/210
0/210
0/210
Appendix
Table 8 : Provider-specific benign false positives under post hoc D2 ratio-branch variants.
Policy
Replay
Long tasks
Blocked ( N=41 )
Tasks done ( N=60 )
Pre-comp. D3 stops ( N=60 )
Pre-comp. D4 stops ( N=60 )
Frozen moderate
41/41
0/60
60/60
0/60
Successor + D4
41/41
53/60
0/60
7/60
Successor, D1–D3
41/41
60/60
0/60
0/60
Appendix
Table 9 : Growth-gated D3 outcomes on 41 replay attacks and 60 completed 8–12-step workflows.
Policy
Protocol outcomes ( N=12 )
Tasks complete
Pre-comp. blocks
Infra. failures
Model incomplete
Fixed $0.10 (counterfactual)
7/12
2/12
1/12
2/12
Checkpoint budget (live)
9/12
0/12
1/12
2/12
Appendix
Table 10 : D3/D4 successor check on 12 Cerebras GPT-OSS-120B workflows.
Metric
Fixed
Progress-auth.
ITT task success
13/24 (54.2%)
22/24 (91.7%)
Pre-completion D4 stop
9/24
0/24
Tool failure
2/24
2/24
Mean cost/session (USD)
0.0992
0.1126
Median cost/session (USD)
0.1037
0.1143
Appendix
Table 11 : D4-only Mistral Small 4 transfer: 24 post-qualification workflows independently executed under each policy (48 sessions).
Treatment
Task success
Cost ratio
Effective / benign
Gross / projected
Groq
Persistent rebilling
5/5
1.249 ×
1.675 ×
Structural-loop proxy
5/5
0.880 ×
1.344 ×
Cost-optimization proxy
5/5
1.192 ×
1.593 ×
Matched benign
5/5
1.000 ×
1.555 ×
Appendix
Table 12 : Success-gated cost comparison of related mechanisms on matched workflows.
Figure 10 : D2 operating point across 1,050 benign provider–scenario executions. Left: 2,292 post-seed turns from 984 multi-turn scenarios; three raw-trigger rows belong to two scenarios, and Δp/p1 uses a logarithmic axis. Right: false-positive counts retain the 1,050-scenario denominator, while attack leakage averages trigger-inclusive fractions over six replay attacks.
Figure 11 : Live inline placement across five attack vectors and five provider–model pairs. Fill encodes first block turn; text gives interception layer/turn; hatching marks an upstream HTTP 413 cap. D2/D3 intercept 23/25 sessions before the next provider call; the other two stop at the provider cap.
Figure 12 : Evaluation pipeline from isolated execution to offline accounting. Each fresh API session emits one execution record to the append-only ledger, from which offline accounting builds the analysis view used for reported results.
Figure 13 : Attack-vector tradeoffs for V0, V1, V3, and V5 ( n=50/36/46/39 executions). Panels report mean CAF, mean turns, and CAF stability (inverse mean within-family CV across six family cells); higher stability is better. No vector leads on all three, and each panel retains its native scale.
Figure 14 : Prevalence of code-visible safeguard proxies across 3,830 public MCP server and transport repositories. (a) The scan detects no proxy in 3,759/3,830 repositories (98.1%) and at least one in 71/3,830 (1.9%). (b) Among those 71 positives, 67 (94.4%) match one family, 4 (5.6%) match two or three, and none match all four.
Tool-augmented LLM agents can harbor implicit state that persists across sessions, activates through events, and propagates across agent boundaries---largely invisible to standard debugging. We formalize this as Persistent Semantic Entities (PSEs): constructs defined by name binding, event triggering, and cross-boundary propagation, and evaluate them across 24 models from 11 families (1.5B--1T parameters). First, every tested model is susceptible (20--100% on the 20-model susceptibility panel), with name binding as the necessary and dominant mechanism: without it, contamination is 0%. Second, persistence depends on contamination type rather than scale or deployment: preference contamination persists undecayed on every model probed (100% at t=10) and instruction contamination persists wherever adopted, persona-style injection decays partially (90%→10%), while factual injection is model-dependent---self-corrected on Llama-3.1-8B and GPT-4o-mini but held at ceiling on both Qwen2.5-coder variants, so we do not claim it self-corrects in general. The preference and instruction results hold across providers in our controlled setting. Third, context-isolated self-verification achieves 20--79% reduction (median 36.5%) without oracle references while keyword-based detection produces systematic false positives, and contamination compounds 1.9× along a four-stage agent pipeline (40%→75%). Preference and instruction contamination---persistent, lacking self-correction, and poorly captured by standard monitoring---represent a particularly concerning attack surface for deployed agent systems.
Zhaohui Wang
USC Viterbi School of Engineering, University of Southern California, Los Angeles, CA, USA.
Persistent memory in LLM agents creates an attack surface that production safety classifiers do not observe: the payload enters via RAG retrieval and persists across sessions via tool-mediated memory. We evaluate six defenses across four architectural layers against delayed-trigger attacks on nine open-source models (5,040 runs, N=40 per condition). Five of six defenses fail: input-level filters never see the payload (it enters via RAG, not user input); retrieval-level classifiers observe it but cannot distinguish compliance-framed injection from legitimate policy; instruction-level hardening is overridden by the stored rule's compliance framing. Only tool-gating at the memory layer (Memory Sandbox) reduces ASR to 0% for eight of nine models, with zero utility cost. A reasoning model inverts this defense via goal-directed RAG fallback, a mechanism that replicates cross-family on Bedrock. A reasoning-mode ablation reveals a double dissociation: no single sandbox implementation is safe across both reasoning and non-reasoning model classes. We resolve this with a content-layer proof-of-concept (RATG), validated on non-reasoning models. A loaded-corpus frontier evaluation (21 models, 3 providers, N=40) overturns an initial empty-corpus screen showing 0/210 exfiltrations: that was a threat-model artifact, not model safety. Under realistic conditions, Gemini 3.1 Pro Preview exfiltrates at 95% ASR, GPT-5.1 regresses to 22.5% relative to GPT-5 (5%), and Anthropic blocks at the injection layer (0-17.5% storage, 0% ASR). Nearly all OpenAI and Gemini models store the rule at 100% regardless of execution resistance, creating supply-chain risk in shared-memory deployments. Defense effectiveness is determined by architectural layer and reasoning capability, not classifier quality.
We introduce MemLineage, a defense for LLM agent memory that attaches both cryptographic provenance and LLM-mediated derivation lineage to every entry. Recent and concurrent work shows that untrusted content can be written into persistent agent state and re-enter later sessions as an instruction; the remaining systems question is how to preserve useful memory recall while preventing such state from justifying sensitive actions. MemLineage treats this as a chain-of-custody problem rather than a filtering problem. It is a six-module design around an RFC-6962 Merkle log over per-principal Ed25519-signed entries: a weighted derivation DAG records which retrieved entries influenced each new memory, and a max-of-strong-edges propagation rule makes Untrusted-Path Persistence hold for any chain whose attribution edges remain above threshold. The sensitive-action gate then refuses dispatches whose active justification descends from an external ancestor, while still allowing benign recall. We evaluate three defense cells against three memory-poisoning workloads on a deterministic mechanism-isolation harness; MemLineage is the only configuration in that harness that drives all three columns to zero ASR, while sub-millisecond per-operation overhead keeps it well below the noise floor of any LLM call. A Codex-backed AgentDojo bridge further separates strong-model behavior from defense-layer behavior: under an intentionally vulnerable tool-output profile, no-defense and signature-only baselines fail on all six banking pairs, while all MemLineage rows reduce strict AgentDojo ASR to zero. The core deterministic artifacts are byte-equal CI-verified; hosted-model AgentDojo and live-model sweeps are recorded as auditable logs rather than byte-pinned artifacts.
Ciyan Ouyang, Rui Hou
State Key Laboratory of Cyberspace Security Defense Institute of Information Engineering, CAS Beijing, China