cs.LGSep 30, 2026

Representation Transitions Reveal Emerging Safety Risks in Multi-Turn LLM Agents

Authors: Haoyu Wang, Wei Zhao, Yedi Zhang, Christopher M. Poskitt, Jun Sun

Organizations: Singapore Management University · National University of Singapore

Abstract

Multi-turn attacks on agentic systems can compose individually permissible actions into harmful outcomes, challenging defenses that assess actions or states in isolation. We show that such attacks leave a detectable signature in the agent's internal representations: harmful behavior emerges as an accumulated representation transition across context updates, whose triggering context can be identified from the same signal. We further find that naive aggregation is confounded by benign representation drift, as a contrastive safety direction need not assign zero to benign transitions. We address this by denoising the direction, anchoring benign traffic at zero and removing its leading variation directions, with no runtime cost. These findings motivate DART, a runtime framework that detects and attributes representation shifts and intervenes with targeted reminders. Across six models and two multi-turn benchmarks, DART reduces attack success from 84% to 25% on MT-AgentRisk, catching every attack at a mean false-alarm rate of 12%, and from 97% to 52% on ASEval, at costs in benign non-refusal of 8% and 0%, respectively. On MT-AgentRisk, it outperforms ToolShield, the state-of-the-art multi-turn defense, on all six models: under the same protocol, ToolShield reaches only 55%. Denoising is critical: on ASEval, the undenoised monitor catches only 7%-40% of attacks, while the denoised monitor catches 60%-85%. The same monitor covers single-turn indirect injection without modification and adds only 0.14-0.56 s overhead per monitored step without requiring an auxiliary model, making it a lightweight complement to computation-heavy speculative defenses.

Figures & tables

Appendix figures & tables18 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. One Turn Too Late: Learning When to Intervene Against Multi-Turn Malicious Intent

    May 7, 2026Xinjie Shen, Rongzhe Wei, Peizhi Niu +6AttackerMalicious Agents

  2. Speculative Safety Honeypot: Toward Proactive Defense Against Multi-turn Agent Attacks

    Sep 30, 2026Zezhong Wang, Xueyang Tang, Rui Lian +2HoneypotsArtificial Intelligence Agents

  3. AMT-X: Phase-Structured Multi-Turn Red-Teaming with Checklist-Gated Evaluation

    Jul 13, 2026Yi Ting Shen, Kentaroh Toyoda, Alex LeungRed-TeamingAttack-Success Rate