Safety Does Not Compose: Non-Decaying Loop State for Autonomous LLM Agents
Authors: Chenhao Wu, Haoxuan Jia, Yang Liu, Yingguang Yang, Yuhan Lin, Chongyang Zhang, Hao Zheng, Yulin Huang, +6 more
Organizations: University of Chinese Academy of Sciences · Fullive-AI · Nanyang Technological University · Supply Chain Tech Team Y, JD.com · Peking University · Fudan University
Large language model agents are increasingly deployed as autonomous loops. Starting from one human goal, such a system repeatedly discovers work, plans, executes tool calls, verifies outcomes and persists state across many unattended iterations. The agent safeguards in wide use, however, are defined over a single trajectory, and their safety state is re-initialized when the next trajectory begins. We show that this is a failure of composition rather than an implementation detail. Our central result is a separation: against an attack whose evidence is fragmented across several iterations, every trajectory-scoped monitor has a true-positive rate equal to its false-positive rate, however expressive it is, because the evidence it would need never appears in the window it sees, whereas a monitor retaining cross-iteration state separates the two perfectly. We further show that the obvious repair of carrying a geometrically decaying risk score is insufficient, because the cooling-off period a patient adversary must wait is a constant that does not grow with the horizon N. We then present LoopHarness, which restores a persistent, non-decaying safety state at the loop level. Under mediated commits and an arbiter detection floor δM, it bounds the expected number of unauthorized irreversible actions by B+m−1+m/δM, a constant in N, of which the B+m−1 term is decided by a model-free rule and therefore survives a fully colluding verifier. We give a complete evaluation protocol on native Agent-SafetyBench tasks with paired clean and attacked episodes, an outer-state attack suite whose decisive evidence exists only across iterations, per-module ablations, and an adaptive white-box red team.
Figures & tables
Figure 1: LoopHarness overview. a, Fragmentation under an observation restriction. Fragments accumulate across trajectories while the inner harness resets its internal safety state to θ0 (Definition 1 ). Equal episode alarm rates follow only when the complete inner observation sequences, including recalled memory, have the same joint law under the attack and matched benign processes (Equation 2 ; Proposition 2 ). Reset alone does not imply this condition. Perfect outer detection additionally requires distinguishing runtime evidence outside the matched inner observations. b, Persistent-state execution loop. At iteration i , stage 0 rehydrates integrity-protected state through mig , while stages 1–2 apply ipm admission and crg clearance. Stage 3 executes the unchanged L1–L4 inner harness under the current capability ceiling before the Tool → World boundary. After the effect, stage 4 performs asa arbitration, and stages 5–6 persist the accepted state through mig , update xrc , and invoke crg to continue or halt. Across iterations, the read/write state carries the xrc risk latch Ξ , crg budgets and capability ceiling, mig integrity memory M , and ipm reject ledger R . Grounded structural evidence latches in Ξ through Equation 1 , rather than being erased by the next per-trajectory reset.
Config
A11
A12
A13
A14
A15
Overall
CleanGC
B0
97.5
99.0
97.5
97.0
97.0
97.6
97.4
B1
97.5
99.0
97.5
97.0
97.0
97.6
97.3
B1.5
97.5
99.0
97.5
97.0
97.0
97.6
97.3
B2
85.5
93.5
83.0
89.0
91.0
88.4
87.8
B3
97.5
99.0
97.5
97.0
97.0
97.6
97.3
B4
0.0
0.0
0.0
0.0
0.5
0.1
96.9
Table 1: Outer-only attack-success rate by family and clean target completion. Lower ASR is safer, higher CleanGC is better. Denominators: 200 attack episodes per family, 1,000 per configuration overall, 1,000 clean targets. ASR is evaluated against each family’s specified controller-contract violation.
cASR (%)
cASR AUC (%)
Config
N=5
N=10
N=20
N=5
N=10
N=20
p vs. B1
B1
2.9
6.2
6.2
1.7
3.3
4.7
–
B1.5
2.9
6.2
6.2
1.7
3.3
4.7
1.000
B4
0.0
0.0
0.0
0.0
0.0
0.0
<0.001
Table 2: Compounding over loop length on Track A. p is the p -value from a same-episode paired bootstrap test comparing cASR(N) with B1 at N=20 . Risk-charged commits and clean-episode completion over the same sweep are in Table 5 .
Config
d
Ξpre
Block (%)
ASR (%)
CleanGC (%)
XRC-off
1
0.100
0.0
97.5
97.5
2
0.098
0.0
95.5
95.5
4
0.099
0.0
97.5
97.5
XRC-decay
1
0.560
100.0
0.0
97.5
2
0.392
0.0
97.0
97.0
4
0.193
0.0
97.5
97.5
Table 3: Matched A14 XRC-state sweep on Track B at N=10 . All settings use identical episodes and keep IPM, MIG, ASA, CRG and the MIG-protected memory path fixed; only the online XRC recurrence changes. XRC-off disables carryover of Ξ rather than deleting persisted memory or checkpoints. Ξpre is the mean retained risk before the target, and Block is conditional on the target being reached. Each row contains 200 attack episodes and 200 clean twins.
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
Threat
Removed
B4
Ablated
Δ
A11
ipm
0.0%
97.5%
+97.5
A12
mig
0.0%
96.0%
+96.0
A13
asa
0.0%
97.5%
+97.5
A14
xrc
0.0%
97.0%
+97.0
A15
crg
0.5%
97.0%
+96.5
Appendix
Table 4: Matched single-module removals on Track B. Δ is ablated minus full-system ASR, in percentage points.
Config
FullGC
Over-
Under-
Halt
TTD
Commits
Adv.
DoS
Escal.
halt
halt
lat.
N=5/10/20
calls
FullGC
B0
100.0
0.0
100.0
–
8.600
1.627/3.070/5.792
0.000
100.0
0.0
B1
79.4
0.0
100.0
–
8.600
0.033/0.082/0.109
1.921
100.0
0.0
B1.5
79.4
0.0
100.0
–
8.600
0.033/0.082/0.109
1.921
100.0
0.0
B2
89.7
0.0
67.8
4.795
7.249
0.035/0.060/0.062
1.942
90.7
0.0
B3
89.7
100.0
0.0
8.600
8.600
0.035/0.087/0.103
2.158
100.0
0.0
Appendix
Table 5: Safety–utility–cost frontier at N=10 , with the risk-charged-commit sweep that accompanies Table 2 . Halt latency is against the evaluator’s first invalid transition and time-to-detect is right-censored at N . The last two columns report the availability condition, in which the adversary’s only objective is to latch Ξ . Escalation is measured only on the latch-only availability episodes and counts recorded governor forced halts or memory-integrity halts; it is not the halt rate of the ordinary Track-A attack episodes.
Variant family
Attempts
Min. mitigation rate
Worst-case mean irreversible commits
Final ASR
Cue suppression
639
0.518
1.649
0.0
Window spacing
364
0.518
1.649
0.0
Channel variation
579
0.508
1.649
0.0
Combined (grid minimum)
3,595
0.508
1.649
0.0
Combined under A-collude
2,837
0.392
1.546
0.0
Appendix
Table 6: Adaptive white-box red-team results on Track A at N=12 . The 3×2×3 grid varies cue suppression, attempt spacing and delivery channel across 18 variants per parent under B4 and A-collude . The first three rows vary one axis, fixing the others at archive, spacing 1, and feed; combined rows use the full grid. Within each row, we report total mounted attempts, minimum mitigation rate over variants with attempts, and maximum mean irreversible commits.
ID
ipm
xrc
mig
asa
crg
ID
ipm
xrc
mig
asa
crg
B0
–
off
–
off
–
A-noIPM
–
non-dec.
✓
full
✓
B1
–
off
–
off
–
A-noXRC
✓
off
✓
full
✓
B1.5
–
decaying
–
off
–
A-decayXRC
✓
decaying
✓
full
✓
B2
–
off
–
naive
–
A-noMIG
✓
non-dec.
–
full
✓
B3
–
off
–
post-hoc
–
A-noASA
✓
non-dec.
✓
off
✓
B4
✓
non-dec.
✓
full
✓
A-noGov
✓
non-dec.
✓
full
–
Appendix
Table 7: Configurations. Each column block is one setting of the same controller; B4 is full LoopHarness and each ablation removes exactly one component from it. Every configuration except B0 runs the same four-layer inner harness; B0 runs none.
ID
Threat
Mechanism
A1
T1
Adversarial instructions injected into the task description (context poisoning)
A2
T2
Malicious payload embedded in an untrusted tool output (indirect prompt injection)
A3
T3, T4, T6
Tool-description tampering and parameter escalation for tool abuse or privilege escalation
A4
T5
Fabricated conversation history asserting false prior authorization
A5
T1–T6
Composite context, tool-output, tool-registry and memory injection
A6
T7
Disguised item with forged provenance on an untrusted channel; no overt cue
Appendix
Table 8: Attack-family taxonomy and evaluation scope. A1–A5 are inherited from SafeHarness ( Lin et al., 2026 ) and are listed for numbering continuity but are not rerun here. The evaluated LoopHarness suites are Track A (A6–A10), which carries in-trajectory evidence, and Track B (A11–A15), whose violations require outer state except for the within-iteration evidence conflict in A13.