Multi-turn jailbreak attacks have emerged as a critical safety threat to LLMs, as harmful objectives are decomposed across a sequence of apparently benign turns to bypass guardrails. Existing defenses lack the reasoning capacity to identify evolving manipulation patterns, often trading helpfulness for safety by over-refusing benign requests related to sensitive topics. We introduce Trace, a multi-turn defense with trajectory-aware structured reasoning. Before generating each response, the model identifies manipulation cues from the trajectory, evaluates both the benign and adversarial interpretations of user intent, assigns a jailbreak score, and commits to an action: Allow, Caution, or Decline. We curate 4k multi-turn adversarial conversations from five attack frameworks, pair them with 2.4k benign dialogs, and 600 sensitive-but-benign conversations. We train Llama-3.1-8B-Instruct with SFT and GRPO under a multi-component reward that jointly optimizes helpfulness on benign prompts and robustness against jailbreak attempts. Across seven multi-turn attack benchmarks, Trace attains an average attack success rate (ASR) of 14.5% against 31.4% for the strongest baseline and 74.9% for the undefended target, while significantly raising the attacker effort required per successful jailbreak. Trace also balances usability and safety, achieving a 93.3% average compliance on over-refusal benchmarks.
Figures & tables
Figure 1: A multi-turn jailbreak attack (left). Without explicit reasoning over the dialog trajectory, the target complies. TRACE generates a structured reasoning trace, extracting manipulation cues, scoring benign and adversarial hypotheses, and aggregating them into a jailbreak score and a calibrated action, successfully defending against the attack.
jt
Trajectory characterization
Action αt
1
Benign request
Allow
2
Sensitive but benign request
Allow
3
Harmful direction
Caution
4
Disguised request demanding harm
Decline
5
Overtly harmful request
Decline
Table 1: Mapping from jailbreak score jt to action αt .
ASR ( ↓ )
Full-Compliance ( ↑ )
Model
X-Tm
Cresc
Actor
CoA
ICON
FITD
AMA
Avg
PHTest
XSTest
Avg
Llama-3.1-8B-Instruct
90.8
74.2
45.0
98.3
86.7
80.8
48.3
74.9
93.2
92.8
93.0
Self-Reminder-MT
71.7
25.8
11.7
70.8
61.7
40.8
25.2
44.0
59.4
62.8
61.1
LLaMA-Guard-3-MT
89.2
49.1
20.8
94.2
70.0
25.8
36.5
55.1
88.6
92.0
90.3
X-Guard
58.3
28.3
19.2
79.2
65.0
37.5
23.3
44.4
83.3
91.6
87.5
Red-Queen-Guard
30.8
20.8
9.2
45.8
80.0
47.5
28.3
37.5
71.1
86.4
78.8
Table 2: Behavior-level attack success rate (ASR, ↓ ) on multi-turn attacks and full-compliance rate ( ↑ ) on over-refusal benchmarks. Best in bold, second-best underlined. Abbreviations: X-Tm (X-Teaming), Cresc (Crescendo), Actor (ActorAttack), and CoA (Chain-of-Attacks).
Benchmark
Base
SFT
GRPO
ARC-Challenge (25-shot)
81.1
80.7
80.6
BBH (3-shot CoT)
61.5
69.5
68.2
GSM-8K (0-shot)
81.1
79.2
79.6
HellaSwag (10-shot)
80.0
77.8
78.1
MMLU-Pro (5-shot CoT)
43.8
44.8
43.5
Table 3: General-capability accuracy (%) for the base (Llama-3.1-8B-Instruct) and Trace variants.
Model
X-Tm
Cresc
Actor
CoA
ICON
FITD
AMA
Avg
Qwen3-8B (Base)
97.5
89.2
37.5
93.3
100.0
89.9
31.7
77.0
Qwen3-8B ( Trace -GRPO)
19.2
14.2
2.5
25.0
1.7
15.0
19.2
13.8
Table 4: Behavior-level ASR across the seven multi-turn attacks for the base Qwen3-8B and Trace training recipe applied to Qwen3-8B. Trace lowers average ASR from 77.0% to 13.8%, closely matching the 14.5% obtained on Llama-3.1-8B-Instruct (Table 2 ).
Target
ASR@3
ASR@5
ASR@10
Avg. Attempts
Llama-3.1-8B
67.5
81.7
90.8
3.5
+ Trace -GRPO
11.7
21.7
38.3
8.0
Qwen3-8B
68.3
80.8
95.8
3.4
+ Trace -GRPO
5.8
10.0
16.7
9.2
Table 5: Single-turn robustness under AutoDAN-Turbo ( Liu et al., 2024a ) . ASR@ k ( ↓ , %), and Avg. Attempts ( ↑ ) to jailbreak per behavior.
Figure 2: The role of cues in TRACE. (a) Percentage of turns at which TRACE flags at least one cue, by turn number. Red shows multi-turn jailbreak attacks; green shows benign multi-turn conversations. (b) Action distribution by turn under multi-turn attacks.
Figure 3: Mapping Trace actions to dual-hypothesis scores for multi-turn attacks. Each bubble represents a (benign,adversarial) score pair, with size proportional to trajectory frequency.
Defense
X-Tm
Cresc
Actor
CoA
ICON
FITD
AMA
Llama-3.1-8B-Instruct
1.0
1.0
1.0
1.0
1.0
1.0
1.0
Self-Reminder-MT
1.7
3.7
4.1
1.8
2.0
2.9
2.3
LLaMA-Guard-3-MT
1.0
1.7
2.2
1.0
1.7
4.8
1.4
X-Guard
2.0
2.0
2.4
1.6
1.8
3.1
2.4
Red-Queen-Guard
4.3
4.5
5.2
3.9
1.1
2.3
2.0
NBF
1.2
1.5
6.5
1.1
31.3
1.2
1.7
Table 6: Defense Robustness Index (DRI) per defense and attack; best in bold.
Trace
STAIR
Setting
n
Action
Mean
P95
Mean
P95
Benign
3,882
—
251
320
594
1,230
PHTest
3,269
—
348
475
499
1,139
Attack
8,772
Allow
403
531
825
1,300
Caution
640
817
—
—
Decline
687
838
224
590
Table 7: Per-turn reasoning-token count for Trace and STAIR across three settings, reported as mean and 95th percentile. Multi-turn attack rows are stratified by the action; STAIR has no native Caution state.
Appendix figures & tables8 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 4: Annotation pipeline. Adversarial and benign trajectories use the State -only annotator variant; sensitive-but-benign trajectories use the full State + Answer variant.
Figure 5: Conversation level prevalence of manipulation cues across attack types in ground-truth annotations. Each conversation can feature multiple cues.
Defense
Training dialog
Multi-turn?
STAIR
50 k
No
Red-Queen-Guard
11.2 k
Yes
NBF-LLM
4 k
Yes
X-Guard
30 k
Yes
Trace (ours)
4.9 k
Yes
Appendix
Table 8: Training corpus size for each fine-tuned defense and whether the corpus consists of multi-turn jailbreak conversations. Self-Reminder-MT and LLaMA-Guard-3-MT are omitted as they do not fine-tune the underlying model.
Pair
Cues
Ben.
Adv.
JB.
Action
Jaccard
Quadratic-weighted κ
H1 vs. H2
0.71
0.89
0.90
0.90
0.92
H1 vs. LLM
0.74
0.87
0.89
0.86
0.82
H2 vs. LLM
0.69
0.85
0.88
0.84
0.81
Appendix
Table 9: Inter-annotator agreement across the State components. H1 and H2 denote the two human annotators; LLM denotes the Claude-Sonnet-4.5 annotator used to label the training data.
Figure 6: Action distribution within the ambiguous adversarial range ( adv∈[5,6] ), stratified by benign score. (a) Multi-turn jailbreak attacks. (b) Single-turn harmless prompts (PHTest). Each bar shows the proportion of allow (green), caution (yellow), and decline (red) actions among turns at that benign value.
Variant
X-Tm
CoA
PHTest
XSTest
ASR ↓
ASR ↓
refusal ↓
refusal ↓
Base
90.8
98.3
6.8
7.2
Trace -GRPO †
0.8
3.3
36.7
10.4
Trace -GRPO
20.8
21.7
7.0
6.4
Appendix
Table 10: Ablation of the three-axis Rcon judge and the harm-adjacent training data. Trace -GRPO † replaces the three-axis judge with a single harm-score signal (with a rule-based hard-refusal filter on Allow actions) and removes the OR-Bench-derived harm-adjacent data from GRPO training; all other components are held fixed.
Hyperparameter
Value
SFT Stage
Fine-tuning method
LoRA
LoRA rank ( r ) / α
32 / 64
LoRA dropout
0.05
LoRA target
All linear layers
Max sequence length
20,480
Appendix
Table 11: Training hyperparameters for the SFT and GRPO stages.
Multi-turn jailbreak attacks pose a growing threat to LLMs by exploiting conversational dynamics such as gradual escalation and cross-turn coordination. Existing defenses either rely on costly retraining -- often degrading model utility -- or apply single-turn analysis independently at each turn, failing to capture how risk accumulates along interaction trajectories. We observe that safety behavior in multi-turn interaction is trajectory-dependent: dialogue history continuously reshapes the model's conditioning context, making it insufficient to evaluate each turn in isolation. Motivated by this insight, we present THRD, the first training-free framework that explicitly models temporal risk accumulation for multi-turn jailbreak defense. THRD integrates four modules: a Turn-level Risk Assessor (TRA) for instantaneous risk estimation, a Historical Context Analyzer (HCA) for cross-turn intent escalation detection, a Response Evaluator (RE) for identifying facilitative outputs, and a Decision Module that combines these signals through a time-evolving scoring mechanism with attenuation-based modulation and trend-aware adjustment. Experiments against state-of-the-art multi-turn attacks -- including tree-search-based and multi-agent collaborative methods -- across two target models show that THRD reduces ASR to 0.2--4.0% while preserving model utility within 1.5% degradation on MMLU and GSM8K. Ablation studies confirm non-redundant module contributions and stable cross-architecture generalization. Analysis of first rejection triggers reveals that over 70% of multi-turn attacks require Turn~2 or later to detect, validating the necessity of explicit temporal aggregation.
Zhiqing Ma, Zhonghao Xu, Dong Yu +3
1Beijing Language and Culture University · †Corresponding authors
Deploying LLMs in multi-turn dialogues facilitates jailbreak attacks that distribute harmful intent across seemingly benign turns. Recent training-based multi-turn jailbreak methods learn long-horizon attack strategies from interaction feedback, but often rely on coarse trajectory-level outcome signals that broadcast uniformly to every turn. However, we find that turn-level contributions in multi-turn jailbreaking are non-uniform, phase-dependent, and target-specific. Such coarse outcome supervision induces a credit assignment problem, leading to over-rewarding redundant turns in successful trajectories and under-crediting useful intermediate turns in failed ones. To address this, we propose TRACE, a turn-aware credit assignment framework for reinforcement learning (RL)-based multi-turn jailbreaking. For successful trajectories, TRACE estimates turn-level contributions via leave-one-turn-out semantic masking; for failed ones, TRACE assigns penalties based on prompt harmfulness and semantic relevance, with an additional local refusal-aware penalty. Furthermore, we reuse the attack-side credit signal for multi-turn defense alignment. Extensive experiments on open-source and closed-source targets show that TRACE achieves strong overall performance in effectiveness, transferability, and efficiency, yielding about a 25% relative improvement in attack success rate over the strongest RL baseline while also improving the safety-utility balance when reused for defense alignment.
Zhida He, Xiaoyu Wen, Han Qi +7
Shanghai AI Laboratory · Fudan University · Shanghai Jiao Tong University
Safety-aligned large language models (LLMs) often refuse a harmful request but comply once the same goal is spread over several turns. Preference objectives score whole responses to single prompts, so their training loss alone cannot control risk on unseen histories. Our analysis gives sufficient conditions under which suppression at supervised single-turn contexts yields a bound on multi-turn trajectory risk. The bound accounts for coverage, transfer slack, and leakage, and characterizes contraction relative to a base-policy risk budget evaluated on the trained policy's contexts. TRACE (Trajectory Return Attribution and Contrastive Erasure) turns this principle into a token-level objective. On the safe response, each token is weighted by the discounted return of a refusal-attributable advantage. The advantage compares a frozen reference model with its refusal-ablated copy, allowing earlier response tokens to receive credit from later refusal-related evidence. At high-gap positions on rejected responses, TRACE combines the observed token with policy-selected alternatives in the erasure target. A gradient-norm penalty replaces the retain set. Across five open-weight models and seven multi-turn attacks, TRACE gives the lowest attack success rate (ASR) in all 35 model and attack pairs, while the model utility evaluated on MMLU and HellaSwag drop by at most 1.23 points. Source code can be found in the supplemental material.
Fengpeng Li, Kemou Li, Qizhou Wang +3
PRADA Lab, King Abdullah University of Science and Technology · State Key Laboratory of Internet of Things for Smart City, University of Macau · Imperfect Information Learning Team, RIKEN Center for Advanced Intelligence Project +1