Multi-turn jailbreak attacks have emerged as a critical safety threat to LLMs, as harmful objectives are decomposed across a sequence of apparently benign turns to bypass guardrails. Existing defenses lack the reasoning capacity to identify evolving manipulation patterns, often trading helpfulness for safety by over-refusing benign requests related to sensitive topics. We introduce Trace, a multi-turn defense with trajectory-aware structured reasoning. Before generating each response, the model identifies manipulation cues from the trajectory, evaluates both the benign and adversarial interpretations of user intent, assigns a jailbreak score, and commits to an action: Allow, Caution, or Decline. We curate 4k multi-turn adversarial conversations from five attack frameworks, pair them with 2.4k benign dialogs, and 600 sensitive-but-benign conversations. We train Llama-3.1-8B-Instruct with SFT and GRPO under a multi-component reward that jointly optimizes helpfulness on benign prompts and robustness against jailbreak attempts. Across seven multi-turn attack benchmarks, Trace attains an average attack success rate (ASR) of 14.5% against 31.4% for the strongest baseline and 74.9% for the undefended target, while significantly raising the attacker effort required per successful jailbreak. Trace also balances usability and safety, achieving a 93.3% average compliance on over-refusal benchmarks.
Figures & tables
Figure 1: A multi-turn jailbreak attack (left). Without explicit reasoning over the dialog trajectory, the target complies. TRACE generates a structured reasoning trace, extracting manipulation cues, scoring benign and adversarial hypotheses, and aggregating them into a jailbreak score and a calibrated action, successfully defending against the attack.
jt
Trajectory characterization
Action αt
1
Benign request
Allow
2
Sensitive but benign request
Allow
3
Harmful direction
Caution
4
Disguised request demanding harm
Decline
5
Overtly harmful request
Decline
Table 1: Mapping from jailbreak score jt to action αt .
ASR ( ↓ )
Full-Compliance ( ↑ )
Model
X-Tm
Cresc
Actor
CoA
ICON
FITD
AMA
Avg
PHTest
XSTest
Avg
Llama-3.1-8B-Instruct
90.8
74.2
45.0
98.3
86.7
80.8
48.3
74.9
93.2
92.8
93.0
Self-Reminder-MT
71.7
25.8
11.7
70.8
61.7
40.8
25.2
44.0
59.4
62.8
61.1
LLaMA-Guard-3-MT
89.2
49.1
20.8
94.2
70.0
25.8
36.5
55.1
88.6
92.0
90.3
X-Guard
58.3
28.3
19.2
79.2
65.0
37.5
23.3
44.4
83.3
91.6
87.5
Red-Queen-Guard
30.8
20.8
9.2
45.8
80.0
47.5
28.3
37.5
71.1
86.4
78.8
Table 2: Behavior-level attack success rate (ASR, ↓ ) on multi-turn attacks and full-compliance rate ( ↑ ) on over-refusal benchmarks. Best in bold, second-best underlined. Abbreviations: X-Tm (X-Teaming), Cresc (Crescendo), Actor (ActorAttack), and CoA (Chain-of-Attacks).
Benchmark
Base
SFT
GRPO
ARC-Challenge (25-shot)
81.1
80.7
80.6
BBH (3-shot CoT)
61.5
69.5
68.2
GSM-8K (0-shot)
81.1
79.2
79.6
HellaSwag (10-shot)
80.0
77.8
78.1
MMLU-Pro (5-shot CoT)
43.8
44.8
43.5
Table 3: General-capability accuracy (%) for the base (Llama-3.1-8B-Instruct) and Trace variants.
Model
X-Tm
Cresc
Actor
CoA
ICON
FITD
AMA
Avg
Qwen3-8B (Base)
97.5
89.2
37.5
93.3
100.0
89.9
31.7
77.0
Qwen3-8B ( Trace -GRPO)
19.2
14.2
2.5
25.0
1.7
15.0
19.2
13.8
Table 4: Behavior-level ASR across the seven multi-turn attacks for the base Qwen3-8B and Trace training recipe applied to Qwen3-8B. Trace lowers average ASR from 77.0% to 13.8%, closely matching the 14.5% obtained on Llama-3.1-8B-Instruct (Table 2 ).
Target
ASR@3
ASR@5
ASR@10
Avg. Attempts
Llama-3.1-8B
67.5
81.7
90.8
3.5
+ Trace -GRPO
11.7
21.7
38.3
8.0
Qwen3-8B
68.3
80.8
95.8
3.4
+ Trace -GRPO
5.8
10.0
16.7
9.2
Table 5: Single-turn robustness under AutoDAN-Turbo ( Liu et al., 2024a ) . ASR@ k ( ↓ , %), and Avg. Attempts ( ↑ ) to jailbreak per behavior.
Figure 2: The role of cues in TRACE. (a) Percentage of turns at which TRACE flags at least one cue, by turn number. Red shows multi-turn jailbreak attacks; green shows benign multi-turn conversations. (b) Action distribution by turn under multi-turn attacks.
Figure 3: Mapping Trace actions to dual-hypothesis scores for multi-turn attacks. Each bubble represents a (benign,adversarial) score pair, with size proportional to trajectory frequency.
Defense
X-Tm
Cresc
Actor
CoA
ICON
FITD
AMA
Llama-3.1-8B-Instruct
1.0
1.0
1.0
1.0
1.0
1.0
1.0
Self-Reminder-MT
1.7
3.7
4.1
1.8
2.0
2.9
2.3
LLaMA-Guard-3-MT
1.0
1.7
2.2
1.0
1.7
4.8
1.4
X-Guard
2.0
2.0
2.4
1.6
1.8
3.1
2.4
Red-Queen-Guard
4.3
4.5
5.2
3.9
1.1
2.3
2.0
NBF
1.2
1.5
6.5
1.1
31.3
1.2
1.7
Table 6: Defense Robustness Index (DRI) per defense and attack; best in bold.
Trace
STAIR
Setting
n
Action
Mean
P95
Mean
P95
Benign
3,882
—
251
320
594
1,230
PHTest
3,269
—
348
475
499
1,139
Attack
8,772
Allow
403
531
825
1,300
Caution
640
817
—
—
Decline
687
838
224
590
Table 7: Per-turn reasoning-token count for Trace and STAIR across three settings, reported as mean and 95th percentile. Multi-turn attack rows are stratified by the action; STAIR has no native Caution state.
Appendix figures & tables8 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 4: Annotation pipeline. Adversarial and benign trajectories use the State -only annotator variant; sensitive-but-benign trajectories use the full State + Answer variant.
Figure 5: Conversation level prevalence of manipulation cues across attack types in ground-truth annotations. Each conversation can feature multiple cues.
Defense
Training dialog
Multi-turn?
STAIR
50 k
No
Red-Queen-Guard
11.2 k
Yes
NBF-LLM
4 k
Yes
X-Guard
30 k
Yes
Trace (ours)
4.9 k
Yes
Appendix
Table 8: Training corpus size for each fine-tuned defense and whether the corpus consists of multi-turn jailbreak conversations. Self-Reminder-MT and LLaMA-Guard-3-MT are omitted as they do not fine-tune the underlying model.
Pair
Cues
Ben.
Adv.
JB.
Action
Jaccard
Quadratic-weighted κ
H1 vs. H2
0.71
0.89
0.90
0.90
0.92
H1 vs. LLM
0.74
0.87
0.89
0.86
0.82
H2 vs. LLM
0.69
0.85
0.88
0.84
0.81
Appendix
Table 9: Inter-annotator agreement across the State components. H1 and H2 denote the two human annotators; LLM denotes the Claude-Sonnet-4.5 annotator used to label the training data.
Figure 6: Action distribution within the ambiguous adversarial range ( adv∈[5,6] ), stratified by benign score. (a) Multi-turn jailbreak attacks. (b) Single-turn harmless prompts (PHTest). Each bar shows the proportion of allow (green), caution (yellow), and decline (red) actions among turns at that benign value.
Variant
X-Tm
CoA
PHTest
XSTest
ASR ↓
ASR ↓
refusal ↓
refusal ↓
Base
90.8
98.3
6.8
7.2
Trace -GRPO †
0.8
3.3
36.7
10.4
Trace -GRPO
20.8
21.7
7.0
6.4
Appendix
Table 10: Ablation of the three-axis Rcon judge and the harm-adjacent training data. Trace -GRPO † replaces the three-axis judge with a single harm-score signal (with a rule-based hard-refusal filter on Allow actions) and removes the OR-Bench-derived harm-adjacent data from GRPO training; all other components are held fixed.
Hyperparameter
Value
SFT Stage
Fine-tuning method
LoRA
LoRA rank ( r ) / α
32 / 64
LoRA dropout
0.05
LoRA target
All linear layers
Max sequence length
20,480
Appendix
Table 11: Training hyperparameters for the SFT and GRPO stages.
PRADA Lab, King Abdullah University of Science and Technology · State Key Laboratory of Internet of Things for Smart City, University of Macau · Imperfect Information Learning Team, RIKEN Center for Advanced Intelligence Project +1