cs.CRAug 27, 2026

Safety Does Not Compose: Non-Decaying Loop State for Autonomous LLM Agents

Authors: Chenhao Wu, Haoxuan Jia, Yang Liu, Yingguang Yang, Yuhan Lin, Chongyang Zhang, Hao Zheng, Yulin Huang, +6 more

Organizations: University of Chinese Academy of Sciences · Fullive-AI · Nanyang Technological University · Supply Chain Tech Team Y, JD.com · Peking University · Fudan University

Abstract

Large language model agents are increasingly deployed as autonomous loops. Starting from one human goal, such a system repeatedly discovers work, plans, executes tool calls, verifies outcomes and persists state across many unattended iterations. The agent safeguards in wide use, however, are defined over a single trajectory, and their safety state is re-initialized when the next trajectory begins. We show that this is a failure of composition rather than an implementation detail. Our central result is a separation: against an attack whose evidence is fragmented across several iterations, every trajectory-scoped monitor has a true-positive rate equal to its false-positive rate, however expressive it is, because the evidence it would need never appears in the window it sees, whereas a monitor retaining cross-iteration state separates the two perfectly. We further show that the obvious repair of carrying a geometrically decaying risk score is insufficient, because the cooling-off period a patient adversary must wait is a constant that does not grow with the horizon NN. We then present LoopHarness, which restores a persistent, non-decaying safety state at the loop level. Under mediated commits and an arbiter detection floor δMδ_M, it bounds the expected number of unauthorized irreversible actions by B+m−1+m/δMB+m-1+m/δ_M, a constant in NN, of which the B+m−1B+m-1 term is decided by a model-free rule and therefore survives a fully colluding verifier. We give a complete evaluation protocol on native Agent-SafetyBench tasks with paired clean and attacked episodes, an outer-state attack suite whose decisive evidence exists only across iterations, per-module ablations, and an adaptive white-box red team.

Figures & tables

Appendix figures & tables5 assets

Supplementary material from the paper’s appendix.

Appendix