Multi-turn attacks on agentic systems can compose individually permissible actions into harmful outcomes, challenging defenses that assess actions or states in isolation. We show that such attacks leave a detectable signature in the agent's internal representations: harmful behavior emerges as an accumulated representation transition across context updates, whose triggering context can be identified from the same signal. We further find that naive aggregation is confounded by benign representation drift, as a contrastive safety direction need not assign zero to benign transitions. We address this by denoising the direction, anchoring benign traffic at zero and removing its leading variation directions, with no runtime cost. These findings motivate DART, a runtime framework that detects and attributes representation shifts and intervenes with targeted reminders. Across six models and two multi-turn benchmarks, DART reduces attack success from 84% to 25% on MT-AgentRisk, catching every attack at a mean false-alarm rate of 12%, and from 97% to 52% on ASEval, at costs in benign non-refusal of 8% and 0%, respectively. On MT-AgentRisk, it outperforms ToolShield, the state-of-the-art multi-turn defense, on all six models: under the same protocol, ToolShield reaches only 55%. Denoising is critical: on ASEval, the undenoised monitor catches only 7%-40% of attacks, while the denoised monitor catches 60%-85%. The same monitor covers single-turn indirect injection without modification and adds only 0.14-0.56 s overhead per monitored step without requiring an auxiliary model, making it a lightweight complement to computation-heavy speculative defenses.
Figures & tables
Figure 1: Overview of DART , which detects harmful representation shifts accumulated across turns, attributes them to the most influential context segment, and mitigates them with a targeted reminder without deleting context or halting the agent.
ASEval ( k/T=0.26 )
MT-AgentRisk ( k/T=1 )
Model
AUROC
catch
FPR
AUROC
catch
FPR
Qwen3-8B
0.712 → 0.958
0.33 → 0.78
0.00 → 0.00
1.000 → 1.000
1.00 → 1.00
0.00 → 0.00
Llama-3.1-8B
0.644 → 0.897
0.25 → 0.75
0.00 → 0.00
1.000 → 1.000
1.00 → 1.00
0.12 → 0.05
Qwen3-14B
0.603 → 0.979
0.07 → 0.75
0.12 → 0.00
0.991 → 0.999
1.00 → 1.00
0.15 → 0.05
Qwen3-30B-A3B
0.573 → 0.946
0.12 → 0.85
0.00 → 0.00
1.000 → 1.000
1.00 → 1.00
0.35 → 0.17
Mistral-24B
0.710 → 0.943
0.40 → 0.60
0.00 → 0.00
0.999 → 1.000
1.00 → 1.00
0.28 → 0.28
Table 1: Denoising separates multi-turn attacks from benign trajectories. Plain → denoised on the same held-out trajectories; plain is the same contrast without the projection, so each pair differs only in denoising. k/T is the share of segments carrying signal; diagnostics are in Table 2 .
Figure 2: Choosing the rank, and reading the subspace it removes (ASEval, Qwen3-8B). (a) Trajectory AUROC against rank p ; p=0 is the plain direction, p=1 removes the benign mean alone. (b) Share of each removed direction’s variance explained by a property of the inducing segment ( R2 for length and position, η2 for the user/tool split and tool identity).
Figure 3: Attack success versus benign utility across two benchmarks and six models. Each dot denotes one defense setting: no defense, DART , DART without denoising (the same monitor and reminder applied to the plain direction), and ToolShield ( Li et al., 2026b ) . ToolShield is always on, whereas DART and its ablation operate at a 10% false-alarm budget. Lower right is better. The ToolShield dots use its released experiences: in domain on MT-AgentRisk, and mapped by tool category on ASEval.
Figure 4: Where the running aggregate peaks , relative to the first injected segment, on held-out harmful ASEval trajectories. Without denoising, the aggregate often peaks before the attack arrives; with denoising, the peak concentrates on the injected segment.
Appendix figures & tables18 assets
Supplementary material from the paper’s appendix.
Appendix
benign
harmful trajectory
segment
Bench.
Model
dir.
E[r∣ben]
at t=0
aligned
traj. AUROC
seg. AUROC
attrib.
ASEval
Qwen3-8B
plain
−5.98†
1.00†
0.07
0.409
0.828
0.750
denoised
+1.02†
0.00†
1.00
0.857
0.723
0.800
Qwen3-32B
plain
−28.07†
1.00†
0.07
0.477
0.791
0.700
denoised
+3.25†
0.00†
1.00
0.894
0.766
0.775
MT-AgentRisk
Qwen3-8B
plain
−0.32
0.00
–
1.000
–
–
Appendix
Table 2: The transition signal, before and after denoising. “At t=0 ” is the share of harmful trajectories whose running aggregate peaks on the opening segment; “aligned” is the share whose peak falls at or after the first injected segment; “attrib.” is how often the largest single contribution is an injected segment (chance 0.299 ). Segment and attribution columns need per-segment labels and are defined only on ASEval; a decomposition attack has no individually harmful turn. Denoised rows use p=16 on ASEval and p=2 on MT-AgentRisk. The residual benign drift of the denoised directions is positive: it lifts both classes alike and does not pull the peak to the opening segment. These are not Table 1 ’s numbers. The ASEval block of this table comes from the diagnostic pass on which the mechanism was characterised: the daggered columns are read off the fitting harness without generation, and the remaining ASEval columns were measured on an earlier build of the ASEval pairing rather than on the final 40+40 evaluation split. It is reported to expose the mechanism —the sign of the drift, where the peak lands, and that the segment-level classifier does not move—and the whole block should be read that way, column by column against its own plain/denoised counterpart, never cell by cell against Table 1 . That is why Qwen3-8B reads 0.857 here and 0.958 there. Table 1 is the number that reflects deployment ; the paragraph “The same pattern online” below re-derives these diagnostics on the deployment split for one cell and finds the same pattern. The MT-AgentRisk rows are online.
R2
η2
direction
var%
length
position
tool vs. user
which tool
Qwen3-8B, layer 19
δˉ− (benign mean)
–
0.00
0.57
0.11
0.11
PC1
33.0%
0.06
0.29
0.51
0.51
PC2
19.6%
0.08
0.09
0.46
0.47
PC3
6.3%
0.13
0.24
0.00
0.33
Appendix
Table 3: What the removed directions encode (ASEval benign fitting transitions, n=167 ). Share of each direction’s variation explained by a property of the segment that induced it: R2 for the continuous properties (segment length in tokens, position in the trajectory), η2 for the categorical ones (user turn versus tool response, and tool identity). Values above 0.3 are bold. The top 16 components hold 82% (Qwen3-8B) and 80% (Qwen3-32B) of benign transition variance.
rank p
0 (plain)
1 (mean only)
2
4
8
16
trajectory AUROC
0.712
0.750
0.912
0.929
0.966
0.958
catch
0.33
0.45
0.72
0.75
0.65
0.78
FPR
0.00
0.20
0.00
0.10
0.00
0.00
Appendix
Table 4: Separation against denoising rank (ASEval, Qwen3-8B, evaluation split, online; shown as a diagnostic, not as the selection criterion—see text). p=0 is the plain contrastive direction; p=1 removes the benign mean only; p≥2 additionally removes the p−1 leading eigenvectors of the benign transition covariance. Catch and FPR are at the 10% false-alarm budget. The deployed setting is p=16 , the largest rank at which the denoising condition still holds.
traj. AUROC
catch / FPR
E[r∣ben]
∣cos(δ^ℓ−,u)∣
rank p
deflate
projector
deflate
projector
deflate
projector
deflate
projector
1
0.750
0.750
0.45 / 0.20
0.45 / 0.20
+0.12
+0.12
0.000
0.000
2
0.912
0.911
0.72 / 0.00
0.72 / 0.00
−1.17
−0.17
0.056
0.000
4
0.929
0.930
0.75 / 0.10
0.78 / 0.10
+0.20
+0.54
0.026
0.000
8
0.966
0.925
0.65 / 0.00
0.68 / 0.00
+0.73
+0.55
0.013
0.000
16
0.958
0.965
0.78 / 0.00
0.70 / 0.00
+0.91
+0.52
0.036
0.000
Appendix
Table 5: The two denoising constructions agree (ASEval, Qwen3-8B, evaluation split, online). “deflate” is sequential deflation against the un-orthonormalised basis, which produced the numbers reported elsewhere in this paper; “projector” is u∝(I−QpQp⊤)wℓ with Qp the Gram–Schmidt basis of the same subspace, δ^ℓ− first. Only the projector satisfies the zero-mean anchor exactly at every rank (last pair of columns), and it reduces the residual benign drift accordingly, but detection is unchanged within the resolution of a 40+40 split. The layer is re-selected independently for each construction, which is why p=8 picks different layers ( 19 vs. 21 ) and shows the largest AUROC gap.
authored on ↓ / applied to →
wksp
bank
slack
trav
workspace
0.06
0.00
0.00
0.02
banking
0.00
0.02
0.00
0.00
slack
0.00
0.00
0.40
0.00
travel
0.20
0.00
0.00
0.02
Appendix
Table 6: AgentSpec rules do not transfer across domains. Each entry is the catch rate—fraction of surfaced injections whose violating action fires the rule—with the rule authored on the row suite and applied to the column suite. Non-zero essentially only on the diagonal: a rule written for one suite’s tools and authorized entities is blind to another’s. Contrast the reading direction, fit on three suites and detecting the held-out fourth at 0.85 – 0.96 . (Weak diagonals for workspace / banking / travel reflect low compliance , not rule failure—AgentSpec catches essentially every attack that emits a violating action.)
Model
DART (halt)
AgentSpec
MT
Qwen3-8B
0.00/0.80
1.00/0.80
Qwen3-14B
0.00/0.90
0.90/0.90
Qwen3-32B
0.00/0.90
1.00/0.90
Qwen3-30B-A3B
0.00/0.80
1.00/0.80
Llama-3.1-8B
0.00/1.00
0.80/1.00
Mistral-24B
0.00/0.80
0.80/0.80
Appendix
Table 7: DART at a zero false-alarm budget ( b=0 , i.e. τ set to the maximum benign calibration score—not an absolute threshold of 0 , which the positive residual benign drift of Table 2 would put every trajectory above) vs. the AgentSpec rule baseline. DART here halts on a flag rather than issuing the deployed reminder, so that the comparison isolates coverage —whether the defense fires at all—from the efficacy of the intervention, as ASR / Utility . On MT, AgentSpec is blind—it emits no block (ASR unchanged, 0.80 – 1.00 ) while the monitor closes the attack to 0.00 . On AgentDojo (in-domain) AgentSpec drives ASR to 0 everywhere but, lacking a precision knob, over-blocks benign runs on Llama/Mistral (utility 0.40 / 0.10 , bold); the monitor at b=0 holds utility at its no-intervention level. The conservative budget is what makes halting affordable here: at the deployed b=0.10 a halt costs far more utility (Appendix F.4 ). AgentSpec is the stronger in-domain specialist; the monitor is comparable-or-better where its gate is precise and uniquely covers MT.
detection
injected run: ASR
Benchmark
Model
catch succ
FPR
none
DART
AgentDojo
Qwen3-8B
0.80
0.09
0.18
0.22
slack
Llama-3.1-8B
0.25
0.18
0.07
0.05
Qwen3-14B
0.90
0.18
0.38
0.07
Qwen3-30B-A3B
0.41
0.00
0.31
0.22
Mistral-24B
0.83
0.18
0.55
0.25
Appendix
Table 8: Single-turn indirect injection. The same monitor and the same cumulative aggregate as the multi-turn setting; harm arriving in a single segment contributes its ri to the sum directly. InjecAgent cells are dh (direct harm) / ds (data stealing); benign utility is near-saturated by construction on InjecAgent, whose utility criterion is met before the injection is read. Underlined : the reminder inverting on 8B-class models.
AgentDojo
InjecAgent
+ enhanced
Cross-
Model
(run-lvl)
dh
ds
dh
ds
integ.
Qwen3-8B
0.57
1.00
1.00
1.00
1.00
1.00
Qwen3-14B
0.67
0.67
0.66
0.68
0.66
0.14
Qwen3-32B
0.91
0.99
1.00
1.00
1.00
1.00
Qwen3-30B-A3B (MoE)
0.79
0.98
1.00
1.00
1.00
1.00
Llama-3.1-8B
0.56
0.69
0.89
0.82
0.90
0.78
Appendix
Table 9: Detection generalizes across independent injection benchmarks (run-level AUROC, harmful vs. benign trajectory, n=10 /class). dh is InjecAgent’s direct-harm family and ds its data-stealing family; ds and cross-integration are two-step chained attacks; “ + enhanced” is InjecAgent’s stronger jailbreak-wrapped attack, which holds or improves detection while roughly doubling agent ASR. The AgentDojo column is the harder in-house run-level gate; capable models sit at 0.86 – 1.00 throughout.
detection
without / with DART
Benchmark
Model
catch succ
FPR
benign Util
harmful ASR
MT-AgentRisk
Qwen3-8B
1.00
0.11
0.83 / 0.71
1.00 / 0.40
3.2 turns
Llama-3.1-8B
1.00
0.06
1.00 / 0.94
0.82 / 0.00
Qwen3-14B
1.00
0.11
0.89 / 0.80
0.93 / 0.53
Qwen3-30B-A3B
1.00
0.17
0.91 / 0.77
0.97 / 0.33
Mistral-24B
1.00
0.14
0.97 / 0.86
0.88 / 0.00
Appendix
Table 10: Plain (undenoised) direction on the full multi-turn sets. Causal online monitor on cumulative drift; DART is the attributed reminder. catch succ is the fraction of successful attacks flagged. Outcome columns are without / with mitigation on the same conversations. Underlined : Qwen3-30B-A3B, which fails on ActorBreaker.
Figure 5: Why injection is detectable and intrinsic misuse is not. PCA of transition representations (Qwen3-8B). Left: InjecAgent injected (red) vs. clean (blue) reads form two clusters (best-layer AUROC 1.00 , Fisher ratio 0.53 ). Right: ASB unsafe (red) vs. safe (blue) behavior interleaves (AUROC 0.76 , Fisher 0.21 ). Both AUROCs here are best-layer and in-sample, chosen so the panels are comparable; the held-out cross-validated ASB figure is 0.65 – 0.71 (text). Injection is an external event on a dominant variance axis; misuse is an internal decision leaving a weak, entangled trace.
Model
App.
Fin.
IoT
Prog.
Web
Mean
Qwen3-8B
0.81
0.94
0.65
0.78
0.69
0.77
Qwen3-32B
0.82
0.96
0.61
0.87
0.74
0.80
Appendix
Table 11: R-Judge leave-one-domain-out AUROC (non-injection, transition form) over the benchmark’s five domains: App. (applications and productivity software), Fin. (finance), IoT (connected devices), Prog. (programming and software engineering), and Web (web browsing and search). The monitor generalizes beyond prompt injection; small- n domains (IoT, Web) are noisier.
Model
State hi
Transition δi
Qwen3-8B
0.850
0.917
Qwen3-14B
0.962
0.961
Qwen3-32B
0.886
0.932
Qwen3-30B-A3B
0.900
0.937
Llama-3.1-8B
0.939
0.956
Mistral-Small-24B
0.765
0.847
Appendix
Table 12: Mean leave-one-suite-out AUROC on AgentDojo, held out over its four tool suites ( workspace , banking , slack , travel ), state vs. transition (identical CV, layer 0 excluded). The transition is the stronger default and lifts the weaker models toward the ceiling.
Variant (Qwen3-8B, ℓ=21 )
Score
Flagged
Clean
−14.5
0%
Benign, long (no imperative)
−7.3
0%
Benign imperative (same wrapper)
+8.6
98%
Malicious
+21.5
100%
Appendix
Table 13: The direction fires on any injected imperative, harmful or not (row 3), and is not driven by length (row 2)—primarily a structural, not semantic-safety, signal.
Model
loso
in-dom.
random
reminder
Qwen3-14B
0.10
0.24
0.05
0.52
Qwen3-32B
0.05
0.05
0.05
0.47
Qwen3-30B-A3B (MoE)
0.00
0.07
0.13
0.20
Appendix
Table 14: Clean-flip rate (injected action suppressed and user task preserved) at the operating point, per intervention. The white-box soft suffix ( loso /in-domain) is dominated by a one-line prompt reminder , and does not beat a random -direction suffix on the larger models—the suppression is not coming from the direction.
layer
sep.
α=0
α=0.5
α=1
α=2
frac 0.20
0.3
0.00
0.00
0.00
0.00
frac 0.35
14
0.00
0.00
0.00
0.50
frac 0.50
20
0.00
0.00
0.17
0.67
frac 0.65
61
0.00
0.00
0.00
0.50
frac 0.80
99
0.00
0.00
0.00
0.00
frac 0.90
160
0.00
0.00
0.00
0.00
Appendix
Table 15: Direct-steering clean flip by layer depth (fraction) × steering strength α (Qwen3-14B, pooled slack + travel ). “sep.” is the injected-vs-clean projection separation, which grows ∼500× with depth. Control appears only in a mid-layer, strong- α window; late layers move the projection most yet flip nothing. Cells here are n=6 ; a 6.2× replication ( n=37 , all four suites) reproduces the peak cell at 0.68 and confirms the depth localisation— 0.43 pooled in the frac 0.35 – 0.65 band at α=2 against 0.05 outside and 0.00 unsteered, Fisher p<10−3 .
unattributed
argmaxjrj
top-2 (path)
# discriminating
Model
ASR / Util
ASR / Util
ASR / Util
trajectories
Qwen3-8B
0.300 / 0.825
0.400 / 0.775
0.400 / 0.775
0
Qwen3-14B
0.650 / 0.850
0.525 / 0.875
0.525 / 0.875
0
Qwen3-30B-A3B
0.425 / 0.825
0.325 / 0.825
0.350 / 0.825
2
Qwen3-32B
0.625 / 0.850
0.500 / 0.875
0.500 / 0.875
0
mean
0.500
0.438
0.444
2 / 160
Appendix
Table 16: What the reminder should quote (MT-AgentRisk, 40 harmful and 40 benign trajectories per model, identical direction, identical τ , identical crossing turns). “# discriminating” counts harmful trajectories on which argmax and top- 2 actually name different content; where the crossing occurs at turn 0 the two coincide by construction. Path-based attribution is within noise of the single argmax . Underlined : Qwen3-8B, where quoting a segment is worse than an abstract warning—the same capability gate that inverts the reminder in Table 8 .
ASR / task completion on the flagged runs
Benchmark
Model
n
no interv.
halt
reminder
withhold read
AgentDojo
Qwen3-8B
35
0.23 / 0.31
0.00 / 0.14
0.29 / 0.37
0.00 / 0.29
slack
Llama-3.1-8B
30
0.03 / 0.43
0.00 / 0.00
0.00 / 0.50
0.03 / 0.00
Qwen3-14B
42
0.45 / 0.57
0.00 / 0.12
0.05 / 0.45
0.05 / 0.24
Qwen3-30B-A3B
21
0.33 / 0.62
0.00 / 0.24
0.10 / 0.71
0.05 / 0.71
Mistral-24B
35
0.71 / 0.60
0.00 / 0.14
0.26 / 0.60
0.09 / 0.43
Appendix
Table 17: Fallbacks where the attributed reminder inverts. Measured on the flagged trajectories only ( n per cell), so that every arm acts on the same runs and the comparison isolates the intervention from the detector. “halt” forces a refusal; “withhold read” replaces the attributed segment with a withheld-content placeholder and re-anchors the accumulator. InjecAgent reports ASR only: its utility criterion is satisfied before the injected segment is read, so completion is near-saturated by construction and not comparable to AgentDojo’s. Underlined : cells where the reminder raises ASR above no intervention. Withholding holds ASR ≤0.09 everywhere, including those cells; halting is the most reliable and the most expensive.
Figure 6: Monitoring overhead. Per monitored segment DART performs one prefill and one dot product against u ; bars are normalised to one monitored step, absolute times at right. The monitor reuses the agent’s own weights, so no second model is resident, and the prefill re-reads states the forward pass has already computed—an integrated implementation pays only the dot product, so this is an upper bound. Median over MT-AgentRisk conversations on one M3 Ultra.
Hidden malicious intent in multi-turn dialogue poses a growing threat to deployed large language models (LLMs). Rather than exposing a harmful objective in a single prompt, attackers can distribute their intent across multiple benign-looking turns, making defense a problem not only of whether a dialogue is harmful, but also of when intervention becomes necessary. Existing trace-level labeling approaches provide only coarse safety signals and do not identify this intervention boundary, making it difficult to distinguish timely intervention from premature refusal or a block that comes too late. This work introduces turn-level harm-enabling supervision for multi-turn defense. We define the earliest harm-enabling turn as the first point at which delivering a candidate response would make the accumulated interaction sufficient to enable harmful action. To instantiate this supervision at scale, we construct the Multi-Turn Intent Dataset (MTID), which contains adaptive attack rollouts, matched benign hard negatives, and annotations of this boundary. Using MTID, we train TurnGate, a response-aware monitor that learns when to intervene, and further optimize its policy through multi-turn reinforcement learning. Experiments show that turn-level boundary supervision improves intervention localization, while reinforcement learning further improves the safety--utility trade-off. TurnGate outperforms existing guardrails and multi-turn monitoring baselines, and generalizes across risk domains, attacker pipelines, and target models. Our code is available at https://github.com/Graph-COM/TurnGate.
Xinjie Shen, Rongzhe Wei, Peizhi Niu +6
Georgia Institute of Technology · University of Illinois Urbana-Champaign · UCSD +3
As Large Language Model (LLM) agents are increasingly deployed in complex environments, multi-turn interaction attacks have become a significant security challenge. Existing detection methods typically rely on historical context. However, this retrospective logic struggles to identify deep malicious intents that are split across turns to hide future risks. Inspired by speculative decoding, we propose the Speculative Safety Honeypot (SSH) framework. SSH uses a multi-agent simulation system composed of small LLMs to build an action-level speculate-and-verify workflow. In the speculation stage, SSH predicts future behaviors of the target agent and asynchronously builds a trajectory tree to expose potential risks in advance. In the verification stage, the system uses the target agent's real actions to calibrate and prune the trajectory tree, effectively reducing false positives. As a plug-and-playable component, SSH provides existing detectors with rich decision redundancy beyond the current interaction slice. By judging risk based on the evolution of the entire trajectory tree rather than a single point in time, the system reduces the reliance on the absolute precision of individual detection components. This improves the defense resilience and the warning lead-time of agent systems against complex temporal attacks.
Safety evaluation of large language models (LLMs) relies largely on single-turn attack datasets and single-judge scoring, underestimating risk from adaptive multi-turn adversaries and reporting a single success rate that does not separate partially actionable outputs from those carrying complete operational detail. We propose AMT-X (Adaptive Multi-Turn Exploitation), a phase-structured multi-turn red-teaming framework. Unlike prior multi-turn attacks that rely on ad hoc escalation or free-form per-goal plans, AMT-X casts the attack as an explicit, reproducible multi-phase state machine driven by semantic signals from the victim, and replaces single-judge scoring with a multi-role jury whose phase-conditioned checklists gate success on actionable harm. Across six frontier victim models (queried under their default safety alignment, without added moderation layers) and seven Moderation sub-categories, AMT-X attains overall attack success rates of 97.6-100% under a lenient score threshold, but 66.7-78.6% under a stricter gate requiring complete, real, and operational detail: a gap of up to 33 percentage points between partially and fully actionable harm.