TRACE: Trajectory Return Attribution and Contrastive Erasure for Multi-Turn Safety
Authors: Fengpeng Li, Kemou Li, Qizhou Wang, Haiwei Wu, Jiantao Zhou, Di Wang
Organizations: PRADA Lab, King Abdullah University of Science and Technology · State Key Laboratory of Internet of Things for Smart City, University of Macau · Imperfect Information Learning Team, RIKEN Center for Advanced Intelligence Project · School of Computer Science and Engineering, University of Electronic Science and Technology of China
Safety-aligned large language models (LLMs) often refuse a harmful request but comply once the same goal is spread over several turns. Preference objectives score whole responses to single prompts, so their training loss alone cannot control risk on unseen histories. Our analysis gives sufficient conditions under which suppression at supervised single-turn contexts yields a bound on multi-turn trajectory risk. The bound accounts for coverage, transfer slack, and leakage, and characterizes contraction relative to a base-policy risk budget evaluated on the trained policy's contexts. TRACE (Trajectory Return Attribution and Contrastive Erasure) turns this principle into a token-level objective. On the safe response, each token is weighted by the discounted return of a refusal-attributable advantage. The advantage compares a frozen reference model with its refusal-ablated copy, allowing earlier response tokens to receive credit from later refusal-related evidence. At high-gap positions on rejected responses, TRACE combines the observed token with policy-selected alternatives in the erasure target. A gradient-norm penalty replaces the retain set. Across five open-weight models and seven multi-turn attacks, TRACE gives the lowest attack success rate (ASR) in all 35 model and attack pairs, while the model utility evaluated on MMLU and HellaSwag drop by at most 1.23 points. Source code can be found in the supplemental material.
Figure 2: Overview of TRACE . (I) TAW weights preferred tokens by the discounted return of a refusal-attributable advantage. (II) RLPE constructs proxy-augmented erasure targets at selected rejected positions. (III) A gradient-norm penalty discourages drift without a retain set.
Model
Method
Multi-Turn Attacks
Single-Turn Attacks
Model Utility
ActorAttack
Crescendo
MHJ
SafeDialBench
STAR
RACE
X-Teaming
GCG
AutoDAN
MMLU
HellaSwag
ASR(%) ↓
ASR(%) ↓
ASR(%) ↓
ASR(%) ↓
ASR(%) ↓
ASR(%) ↓
ASR(%) ↓
ASR(%) ↓
ASR(%) ↓
Accuracy(%) ↑
Accuracy(%) ↑
Gemma-3-1B
Original Model
54.00
50.70
63.10
63.30
74.00
42.00
88.70
18.00
19.50
38.70
57.23
SFT
49.86
47.06
61.77
60.34
70.40
38.64
82.50
17.34
13.63
36.83
56.23
DPO
48.08
49.69
57.52
58.07
73.65
41.55
84.27
15.76
12.50
38.52
56.87
B-DPO
48.64
50.35
57.08
58.42
72.94
41.66
83.01
15.70
11.85
38.66
56.52
Table 1: ASR (%) under seven multi-turn and two single-turn attacks, and MMLU and HellaSwag accuracy (%), on Gemma-3-1B, Llama-3.1-8B and Qwen-3.5-9B. Original Model is the released instruction-tuned checkpoint before our fine-tuning. Lower ASR and higher accuracy are better.
Method
Multi-Turn Attacks
Single-Turn Attacks
Model Utility
ActorAttack
Crescendo
MHJ
SafeDialBench
STAR
RACE
X-Teaming
GCG
AutoDAN
MMLU
HellaSwag
ASR ↓
ASR ↓
ASR ↓
ASR ↓
ASR ↓
ASR ↓
ASR ↓
ASR ↓
ASR ↓
Accuracy ↑
Accuracy ↑
Original Model
58.17
56.00
67.97
43.35
61.33
40.44
91.82
8.04
1.25
84.48
83.40
SFT
51.82
50.71
51.43
31.36
55.82
35.05
82.63
6.62
0.020
83.86
83.32
DPO
56.53
52.10
61.56
38.55
56.45
36.46
87.88
7.85
0.048
84.41
83.45
B-DPO
55.31
52.65
60.92
37.38
55.64
36.34
89.75
7.08
0.048
84.57
83.41
Table 2: Results on Qwen-3.5-27B trained with LoRA, in the format of Table 1 .
Method
Multi-Turn Attacks
Single-Turn Attacks
Model Utility
Training Time on H200
ActorAttack
Crescendo
MHJ
SafeDialBench
STAR
RACE
X-Teaming
GCG
AutoDAN
MMLU
HellaSwag
ASR ↓
ASR ↓
ASR ↓
ASR ↓
ASR ↓
ASR ↓
ASR ↓
ASR ↓
ASR ↓
Accuracy ↑
Accuracy ↑
Hours
Original Model
58.17
56.00
67.97
43.35
61.33
40.44
91.82
18.59
2.50
67.81
79.12
SFT
56.00
62.67
70.39
40.21
66.00
39.33
87.42
25.91
1.52
67.52
78.65
1.13
DPO
54.83
64.67
62.01
42.17
60.67
35.56
92.45
15.36
6.81
67.23
78.83
1.53
TRACE w/o- LTAW
50.50
49.33
59.96
35.05
55.33
31.33
71.07
11.81
0.50
67.26
78.61
2.95
Table 3: Ablations on Llama-3.1-8B. Each TRACE row removes one component, except “w/ Retention Dataset”, which replaces the gradient-norm penalty with a W-DOOR-style retain loss. “DPO w/ LTAW ” adds TAW to DPO. TAW denotes trajectory-aware adaptive weighting, RLPE denotes risk-localized proxy erasure, and GNP denotes the gradient-norm penalty. The retain-loss variant uses 400 Alpaca examples in addition to the shared preference tuples. The other TRACE variants use no retain dataset. Each attack column uses the same test items across all rows.
Appendix figures & tables11 assets
Supplementary material from the paper’s appendix.
Appendix
Notation
Description
Models and preference data (Sections 1 , 4.1 )
πθ
Trainable policy with parameters θ
πref=πbase
Frozen initial checkpoint πθ0
πabl,πnew
Refusal-ablated reference; policy after training
Dpref
Preference tuples (x,y+,y−) : request, preferred safe and rejected responses
M+,M−,N
Supervised token positions; N=∣M+∣
Appendix
Table 4: Main notations and their descriptions used in the paper.
Algorithm 1 TRACE training with a fixed surrogate per update
Setting
Full fine-tuning
LoRA
Models
Gemma-3-1B, Llama-3.1-8B, Qwen-3.5-9B
Qwen-3.5-27B, Qwen-3.6-27B
Trainable parameters
all
rank-16 adapters, α=32
Precision
bf16
4-bit base, 8-bit scoring models, bf16 compute
Batch size and accumulation
2 and 1
2 and 1
Training length
2,000 steps (10 epochs)
400 steps (2 epochs)
Learning rate
1×10−5
1×10−5
Appendix
Table 5: Optimization settings of the two fine-tuning regimes.
Base model
Refusal-ablated scorer (Hugging Face repository)
Llama-3.1-8B-Instruct
mlabonne/Meta-Llama-3.1-8B-Instruct-abliterated
Qwen3.5-9B
huihui-ai/Huihui-Qwen3.5-9B-abliterated
Qwen3.5-27B
huihui-ai/Huihui-Qwen3.5-27B-abliterated
Qwen3.6-27B
huihui-ai/Huihui-Qwen3.6-27B-abliterated
Gemma-3-1B-IT
mlabonne/gemma-3-1b-it-abliterated-v2
Appendix
Table 6: Frozen refusal-ablated scoring checkpoints. Repository identifiers specify the public releases used for πabl . The corresponding released instruction-tuned model is πref .
Symbol
Role
Full FT
LoRA
Swept values
λ
discount of the return
0.5
0.5
Table 10
ψ
amplitude of wtraj
1.0
1.0
χ
temperature of wtraj
5.0
1.0
1.0, 2.0, 5.0
Amax
clip of A^j
3.0
3.0
ν
strength of wcap
0.5
0.5
kproxy
position and proxy budget (RLPE)
10
10
Table 11
Appendix
Table 7: TRACE hyperparameters. Swept values refer to Llama-3.1-8B under full fine-tuning.
Protocol
Items
Goals
Generation
Turn budget
Turns observed
ActorAttack
600
dataset-native
live, adaptive
from the dataset
6 / 5.91 / 6
Crescendo
150
HarmBench 100, AdvBench 50
live, adaptive
10 rounds
10 / 10.00 / 10
MHJ
537
dataset-native
replayed
from the dataset
4 / 5.35 / 27
SafeDialBench
2,037
dataset-native
replayed
from the dataset
5 / 4.92 / 9
STAR
150
HarmBench 50, JailbreakBench 100
live, adaptive
7 turns, 3 backtracks
7 / 7.00 / 7
RACE
450
AdvBench 50, HarmBench 400
live, adaptive
3 states, 3 seeds, 3 rounds
3 / 2.94 / 3
Appendix
Table 8: Multi-turn evaluation protocols. Turns are the median, mean and maximum per item.
Method
Multi-Turn
Single-Turn Attacks
Model Utility
ActorAttack
Crescendo
MHJ
SafeDialBench
STAR
RACE
X-Teaming
GCG
AutoDAN
MMLU
HellaSwag
ASR ↓
ASR ↓
ASR ↓
ASR ↓
ASR ↓
ASR ↓
ASR ↓
ASR ↓
ASR ↓
Accuracy ↑
Accuracy ↑
Original Model
41.17
48.67
52.16
17.87
46.67
32.22
91.19
7.04
1.25
84.50
83.41
SFT
63.17
58.67
56.98
40.99
61.33
46.44
87.11
6.03
2.50
84.17
83.08
DPO
65.67
63.33
56.05
40.94
56.67
46.89
78.62
2.79
1.00
84.41
83.56
B-DPO
59.83
60.00
53.82
43.99
55.33
49.11
82.39
3.02
1.50
84.39
83.41
Appendix
Table 9: Results on Qwen-3.6-27B trained with LoRA, in the format of Table 1 .
Model
λ Value
Multi-Turn Attacks
Single-Turn Attacks
Model Utility
ActorAttack
Crescendo
MHJ
SafeDialBench
STAR
RACE
X-Teaming
GCG
AutoDAN
MMLU
HellaSwag
ASR ↓
ASR ↓
ASR ↓
ASR ↓
ASR ↓
ASR ↓
ASR ↓
ASR ↓
ASR ↓
Accuracy ↑
Accuracy ↑
Llama-3.1-8B
0.0
39.00
62.00
28.12
18.23
61.33
36.44
59.75
6.53
1.75
67.26
78.91
0.1
34.77
52.00
24.02
17.43
56.67
34.44
55.35
5.03
1.00
67.06
78.68
0.3
33.00
47.33
23.09
17.57
50.67
32.00
49.68
3.77
0.0
66.94
78.43
0.5
22.14
31.33
22.76
17.57
43.33
18.89
33.96
3.27
0.0
66.71
78.27
Appendix
Table 10: Effect of the discount λ on Llama-3.1-8B, with kproxy=10 and ξ=0.05 .
Model
kproxy Value
Multi-Turn Attacks
Single-Turn Attacks
Model Utility
ActorAttack
Crescendo
MHJ
SafeDialBench
STAR
RACE
X-Teaming
GCG
AutoDAN
MMLU
HellaSwag
ASR ↓
ASR ↓
ASR ↓
ASR ↓
ASR ↓
ASR ↓
ASR ↓
ASR ↓
ASR ↓
Accuracy ↑
Accuracy ↑
Llama-3.1-8B
0
28.93
45.33
27.75
17.43
56.00
31.56
63.52
5.28
2.01
67.25
78.08
5
22.53
32.00
29.24
17.77
60.67
60.67
40.25
4.52
1.00
66.82
78.47
10
22.14
31.33
22.76
17.57
43.33
18.89
33.96
3.27
0.0
66.71
78.27
15
18.71
37.33
29.42
20.08
65.33
65.33
38.36
4.77
1.50
66.99
78.36
Appendix
Table 11: Effect of the shared position and proxy budget kproxy on Llama-3.1-8B, with λ=0.5 and ξ=0.05 .
Model
ξ Value
Multi-Turn Attacks
Single-Turn Attacks
Model Utility
ActorAttack
Crescendo
MHJ
SafeDialBench
STAR
RACE
X-Teaming
GCG
AutoDAN
MMLU
HellaSwag
ASR ↓
ASR ↓
ASR ↓
ASR ↓
ASR ↓
ASR ↓
ASR ↓
ASR ↓
ASR ↓
Accuracy ↑
Accuracy ↑
Llama-3.1-8B
0.0
15.17
24.67
19.37
13.45
37.33
11.11
25.16
0.0
0.0
65.42
75.36
0.01
25.17
39.33
25.70
18.56
48.67
20.22
41.59
5.03
2.50
67.63
78.98
0.02
23.33
36.00
24.21
17.92
46.67
19.56
38.36
4.27
1.00
67.56
78.80
0.05
22.14
31.33
22.76
17.57
43.33
18.89
33.96
3.27
0.0
66.71
78.27
Appendix
Table 12: Effect of the penalty radius ξ on Llama-3.1-8B, with λ=0.5 and kproxy=10 .
Method
Multi-Turn Attacks
Single-Turn Attacks
Model Utility
Over-Refusal
ActorAttack
Crescendo
MHJ
SafeDialBench
STAR
GCG
AutoDAN
MMLU
HellaSwag
XSTest
ASR ↓
ASR ↓
ASR ↓
ASR ↓
ASR ↓
ASR ↓
ASR ↓
Accuracy ↑
Accuracy ↑
Refusal Rate ↓
Original Model
58.17
56.00
67.97
43.35
61.33
18.59
2.50
67.81
79.12
5.48
SFT
56.00
62.67
70.39
40.21
66.00
25.91
1.52
67.52
78.65
11.22
DPO
54.83
64.67
62.01
42.17
60.67
15.36
6.81
67.23
78.83
10.16
B-DPO
54.60
51.90
65.80
47.00
59.33
15.70
7.09
67.45
78.58
10.85
Appendix
Table 13: Over-refusal on Llama-3.1-8B, measured as the refusal rate (%) on the safe prompts of XSTest. The attack and accuracy columns repeat the matching entries of Table 1 and Table 12 .
Multi-turn jailbreak attacks have emerged as a critical safety threat to LLMs, as harmful objectives are decomposed across a sequence of apparently benign turns to bypass guardrails. Existing defenses lack the reasoning capacity to identify evolving manipulation patterns, often trading helpfulness for safety by over-refusing benign requests related to sensitive topics. We introduce Trace, a multi-turn defense with trajectory-aware structured reasoning. Before generating each response, the model identifies manipulation cues from the trajectory, evaluates both the benign and adversarial interpretations of user intent, assigns a jailbreak score, and commits to an action: Allow, Caution, or Decline. We curate 4k multi-turn adversarial conversations from five attack frameworks, pair them with 2.4k benign dialogs, and 600 sensitive-but-benign conversations. We train Llama-3.1-8B-Instruct with SFT and GRPO under a multi-component reward that jointly optimizes helpfulness on benign prompts and robustness against jailbreak attempts. Across seven multi-turn attack benchmarks, Trace attains an average attack success rate (ASR) of 14.5% against 31.4% for the strongest baseline and 74.9% for the undefended target, while significantly raising the attacker effort required per successful jailbreak. Trace also balances usability and safety, achieving a 93.3% average compliance on over-refusal benchmarks.
Multi-turn jailbreak attacks pose a growing threat to LLMs by exploiting conversational dynamics such as gradual escalation and cross-turn coordination. Existing defenses either rely on costly retraining -- often degrading model utility -- or apply single-turn analysis independently at each turn, failing to capture how risk accumulates along interaction trajectories. We observe that safety behavior in multi-turn interaction is trajectory-dependent: dialogue history continuously reshapes the model's conditioning context, making it insufficient to evaluate each turn in isolation. Motivated by this insight, we present THRD, the first training-free framework that explicitly models temporal risk accumulation for multi-turn jailbreak defense. THRD integrates four modules: a Turn-level Risk Assessor (TRA) for instantaneous risk estimation, a Historical Context Analyzer (HCA) for cross-turn intent escalation detection, a Response Evaluator (RE) for identifying facilitative outputs, and a Decision Module that combines these signals through a time-evolving scoring mechanism with attenuation-based modulation and trend-aware adjustment. Experiments against state-of-the-art multi-turn attacks -- including tree-search-based and multi-agent collaborative methods -- across two target models show that THRD reduces ASR to 0.2--4.0% while preserving model utility within 1.5% degradation on MMLU and GSM8K. Ablation studies confirm non-redundant module contributions and stable cross-architecture generalization. Analysis of first rejection triggers reveals that over 70% of multi-turn attacks require Turn~2 or later to detect, validating the necessity of explicit temporal aggregation.
Zhiqing Ma, Zhonghao Xu, Dong Yu +3
1Beijing Language and Culture University · †Corresponding authors
Hidden malicious intent in multi-turn dialogue poses a growing threat to deployed large language models (LLMs). Rather than exposing a harmful objective in a single prompt, attackers can distribute their intent across multiple benign-looking turns, making defense a problem not only of whether a dialogue is harmful, but also of when intervention becomes necessary. Existing trace-level labeling approaches provide only coarse safety signals and do not identify this intervention boundary, making it difficult to distinguish timely intervention from premature refusal or a block that comes too late. This work introduces turn-level harm-enabling supervision for multi-turn defense. We define the earliest harm-enabling turn as the first point at which delivering a candidate response would make the accumulated interaction sufficient to enable harmful action. To instantiate this supervision at scale, we construct the Multi-Turn Intent Dataset (MTID), which contains adaptive attack rollouts, matched benign hard negatives, and annotations of this boundary. Using MTID, we train TurnGate, a response-aware monitor that learns when to intervene, and further optimize its policy through multi-turn reinforcement learning. Experiments show that turn-level boundary supervision improves intervention localization, while reinforcement learning further improves the safety--utility trade-off. TurnGate outperforms existing guardrails and multi-turn monitoring baselines, and generalizes across risk domains, attacker pipelines, and target models. Our code is available at https://github.com/Graph-COM/TurnGate.
Xinjie Shen, Rongzhe Wei, Peizhi Niu +6
Georgia Institute of Technology · University of Illinois Urbana-Champaign · UCSD +3