TRACE: Trajectory Return Attribution and Contrastive Erasure for Multi-Turn Safety
Authors: Fengpeng Li, Kemou Li, Qizhou Wang, Haiwei Wu, Jiantao Zhou, Di Wang
Organizations: PRADA Lab, King Abdullah University of Science and Technology · State Key Laboratory of Internet of Things for Smart City, University of Macau · Imperfect Information Learning Team, RIKEN Center for Advanced Intelligence Project · School of Computer Science and Engineering, University of Electronic Science and Technology of China
Safety-aligned large language models (LLMs) often refuse a harmful request but comply once the same goal is spread over several turns. Preference objectives score whole responses to single prompts, so their training loss alone cannot control risk on unseen histories. Our analysis gives sufficient conditions under which suppression at supervised single-turn contexts yields a bound on multi-turn trajectory risk. The bound accounts for coverage, transfer slack, and leakage, and characterizes contraction relative to a base-policy risk budget evaluated on the trained policy's contexts. TRACE (Trajectory Return Attribution and Contrastive Erasure) turns this principle into a token-level objective. On the safe response, each token is weighted by the discounted return of a refusal-attributable advantage. The advantage compares a frozen reference model with its refusal-ablated copy, allowing earlier response tokens to receive credit from later refusal-related evidence. At high-gap positions on rejected responses, TRACE combines the observed token with policy-selected alternatives in the erasure target. A gradient-norm penalty replaces the retain set. Across five open-weight models and seven multi-turn attacks, TRACE gives the lowest attack success rate (ASR) in all 35 model and attack pairs, while the model utility evaluated on MMLU and HellaSwag drop by at most 1.23 points. Source code can be found in the supplemental material.
Figure 2: Overview of TRACE . (I) TAW weights preferred tokens by the discounted return of a refusal-attributable advantage. (II) RLPE constructs proxy-augmented erasure targets at selected rejected positions. (III) A gradient-norm penalty discourages drift without a retain set.
Model
Method
Multi-Turn Attacks
Single-Turn Attacks
Model Utility
ActorAttack
Crescendo
MHJ
SafeDialBench
STAR
RACE
X-Teaming
GCG
AutoDAN
MMLU
HellaSwag
ASR(%) ↓
ASR(%) ↓
ASR(%) ↓
ASR(%) ↓
ASR(%) ↓
ASR(%) ↓
ASR(%) ↓
ASR(%) ↓
ASR(%) ↓
Accuracy(%) ↑
Accuracy(%) ↑
Gemma-3-1B
Original Model
54.00
50.70
63.10
63.30
74.00
42.00
88.70
18.00
19.50
38.70
57.23
SFT
49.86
47.06
61.77
60.34
70.40
38.64
82.50
17.34
13.63
36.83
56.23
DPO
48.08
49.69
57.52
58.07
73.65
41.55
84.27
15.76
12.50
38.52
56.87
B-DPO
48.64
50.35
57.08
58.42
72.94
41.66
83.01
15.70
11.85
38.66
56.52
Table 1: ASR (%) under seven multi-turn and two single-turn attacks, and MMLU and HellaSwag accuracy (%), on Gemma-3-1B, Llama-3.1-8B and Qwen-3.5-9B. Original Model is the released instruction-tuned checkpoint before our fine-tuning. Lower ASR and higher accuracy are better.
Method
Multi-Turn Attacks
Single-Turn Attacks
Model Utility
ActorAttack
Crescendo
MHJ
SafeDialBench
STAR
RACE
X-Teaming
GCG
AutoDAN
MMLU
HellaSwag
ASR ↓
ASR ↓
ASR ↓
ASR ↓
ASR ↓
ASR ↓
ASR ↓
ASR ↓
ASR ↓
Accuracy ↑
Accuracy ↑
Original Model
58.17
56.00
67.97
43.35
61.33
40.44
91.82
8.04
1.25
84.48
83.40
SFT
51.82
50.71
51.43
31.36
55.82
35.05
82.63
6.62
0.020
83.86
83.32
DPO
56.53
52.10
61.56
38.55
56.45
36.46
87.88
7.85
0.048
84.41
83.45
B-DPO
55.31
52.65
60.92
37.38
55.64
36.34
89.75
7.08
0.048
84.57
83.41
Table 2: Results on Qwen-3.5-27B trained with LoRA, in the format of Table 1 .
Method
Multi-Turn Attacks
Single-Turn Attacks
Model Utility
Training Time on H200
ActorAttack
Crescendo
MHJ
SafeDialBench
STAR
RACE
X-Teaming
GCG
AutoDAN
MMLU
HellaSwag
ASR ↓
ASR ↓
ASR ↓
ASR ↓
ASR ↓
ASR ↓
ASR ↓
ASR ↓
ASR ↓
Accuracy ↑
Accuracy ↑
Hours
Original Model
58.17
56.00
67.97
43.35
61.33
40.44
91.82
18.59
2.50
67.81
79.12
SFT
56.00
62.67
70.39
40.21
66.00
39.33
87.42
25.91
1.52
67.52
78.65
1.13
DPO
54.83
64.67
62.01
42.17
60.67
35.56
92.45
15.36
6.81
67.23
78.83
1.53
TRACE w/o- LTAW
50.50
49.33
59.96
35.05
55.33
31.33
71.07
11.81
0.50
67.26
78.61
2.95
Table 3: Ablations on Llama-3.1-8B. Each TRACE row removes one component, except “w/ Retention Dataset”, which replaces the gradient-norm penalty with a W-DOOR-style retain loss. “DPO w/ LTAW ” adds TAW to DPO. TAW denotes trajectory-aware adaptive weighting, RLPE denotes risk-localized proxy erasure, and GNP denotes the gradient-norm penalty. The retain-loss variant uses 400 Alpaca examples in addition to the shared preference tuples. The other TRACE variants use no retain dataset. Each attack column uses the same test items across all rows.
Appendix figures & tables11 assets
Supplementary material from the paper’s appendix.
Appendix
Notation
Description
Models and preference data (Sections 1 , 4.1 )
πθ
Trainable policy with parameters θ
πref=πbase
Frozen initial checkpoint πθ0
πabl,πnew
Refusal-ablated reference; policy after training
Dpref
Preference tuples (x,y+,y−) : request, preferred safe and rejected responses
M+,M−,N
Supervised token positions; N=∣M+∣
Appendix
Table 4: Main notations and their descriptions used in the paper.
Algorithm 1 TRACE training with a fixed surrogate per update
Setting
Full fine-tuning
LoRA
Models
Gemma-3-1B, Llama-3.1-8B, Qwen-3.5-9B
Qwen-3.5-27B, Qwen-3.6-27B
Trainable parameters
all
rank-16 adapters, α=32
Precision
bf16
4-bit base, 8-bit scoring models, bf16 compute
Batch size and accumulation
2 and 1
2 and 1
Training length
2,000 steps (10 epochs)
400 steps (2 epochs)
Learning rate
1×10−5
1×10−5
Appendix
Table 5: Optimization settings of the two fine-tuning regimes.
Base model
Refusal-ablated scorer (Hugging Face repository)
Llama-3.1-8B-Instruct
mlabonne/Meta-Llama-3.1-8B-Instruct-abliterated
Qwen3.5-9B
huihui-ai/Huihui-Qwen3.5-9B-abliterated
Qwen3.5-27B
huihui-ai/Huihui-Qwen3.5-27B-abliterated
Qwen3.6-27B
huihui-ai/Huihui-Qwen3.6-27B-abliterated
Gemma-3-1B-IT
mlabonne/gemma-3-1b-it-abliterated-v2
Appendix
Table 6: Frozen refusal-ablated scoring checkpoints. Repository identifiers specify the public releases used for πabl . The corresponding released instruction-tuned model is πref .
Symbol
Role
Full FT
LoRA
Swept values
λ
discount of the return
0.5
0.5
Table 10
ψ
amplitude of wtraj
1.0
1.0
χ
temperature of wtraj
5.0
1.0
1.0, 2.0, 5.0
Amax
clip of A^j
3.0
3.0
ν
strength of wcap
0.5
0.5
kproxy
position and proxy budget (RLPE)
10
10
Table 11
Appendix
Table 7: TRACE hyperparameters. Swept values refer to Llama-3.1-8B under full fine-tuning.
Protocol
Items
Goals
Generation
Turn budget
Turns observed
ActorAttack
600
dataset-native
live, adaptive
from the dataset
6 / 5.91 / 6
Crescendo
150
HarmBench 100, AdvBench 50
live, adaptive
10 rounds
10 / 10.00 / 10
MHJ
537
dataset-native
replayed
from the dataset
4 / 5.35 / 27
SafeDialBench
2,037
dataset-native
replayed
from the dataset
5 / 4.92 / 9
STAR
150
HarmBench 50, JailbreakBench 100
live, adaptive
7 turns, 3 backtracks
7 / 7.00 / 7
RACE
450
AdvBench 50, HarmBench 400
live, adaptive
3 states, 3 seeds, 3 rounds
3 / 2.94 / 3
Appendix
Table 8: Multi-turn evaluation protocols. Turns are the median, mean and maximum per item.
Method
Multi-Turn
Single-Turn Attacks
Model Utility
ActorAttack
Crescendo
MHJ
SafeDialBench
STAR
RACE
X-Teaming
GCG
AutoDAN
MMLU
HellaSwag
ASR ↓
ASR ↓
ASR ↓
ASR ↓
ASR ↓
ASR ↓
ASR ↓
ASR ↓
ASR ↓
Accuracy ↑
Accuracy ↑
Original Model
41.17
48.67
52.16
17.87
46.67
32.22
91.19
7.04
1.25
84.50
83.41
SFT
63.17
58.67
56.98
40.99
61.33
46.44
87.11
6.03
2.50
84.17
83.08
DPO
65.67
63.33
56.05
40.94
56.67
46.89
78.62
2.79
1.00
84.41
83.56
B-DPO
59.83
60.00
53.82
43.99
55.33
49.11
82.39
3.02
1.50
84.39
83.41
Appendix
Table 9: Results on Qwen-3.6-27B trained with LoRA, in the format of Table 1 .
Model
λ Value
Multi-Turn Attacks
Single-Turn Attacks
Model Utility
ActorAttack
Crescendo
MHJ
SafeDialBench
STAR
RACE
X-Teaming
GCG
AutoDAN
MMLU
HellaSwag
ASR ↓
ASR ↓
ASR ↓
ASR ↓
ASR ↓
ASR ↓
ASR ↓
ASR ↓
ASR ↓
Accuracy ↑
Accuracy ↑
Llama-3.1-8B
0.0
39.00
62.00
28.12
18.23
61.33
36.44
59.75
6.53
1.75
67.26
78.91
0.1
34.77
52.00
24.02
17.43
56.67
34.44
55.35
5.03
1.00
67.06
78.68
0.3
33.00
47.33
23.09
17.57
50.67
32.00
49.68
3.77
0.0
66.94
78.43
0.5
22.14
31.33
22.76
17.57
43.33
18.89
33.96
3.27
0.0
66.71
78.27
Appendix
Table 10: Effect of the discount λ on Llama-3.1-8B, with kproxy=10 and ξ=0.05 .
Model
kproxy Value
Multi-Turn Attacks
Single-Turn Attacks
Model Utility
ActorAttack
Crescendo
MHJ
SafeDialBench
STAR
RACE
X-Teaming
GCG
AutoDAN
MMLU
HellaSwag
ASR ↓
ASR ↓
ASR ↓
ASR ↓
ASR ↓
ASR ↓
ASR ↓
ASR ↓
ASR ↓
Accuracy ↑
Accuracy ↑
Llama-3.1-8B
0
28.93
45.33
27.75
17.43
56.00
31.56
63.52
5.28
2.01
67.25
78.08
5
22.53
32.00
29.24
17.77
60.67
60.67
40.25
4.52
1.00
66.82
78.47
10
22.14
31.33
22.76
17.57
43.33
18.89
33.96
3.27
0.0
66.71
78.27
15
18.71
37.33
29.42
20.08
65.33
65.33
38.36
4.77
1.50
66.99
78.36
Appendix
Table 11: Effect of the shared position and proxy budget kproxy on Llama-3.1-8B, with λ=0.5 and ξ=0.05 .
Model
ξ Value
Multi-Turn Attacks
Single-Turn Attacks
Model Utility
ActorAttack
Crescendo
MHJ
SafeDialBench
STAR
RACE
X-Teaming
GCG
AutoDAN
MMLU
HellaSwag
ASR ↓
ASR ↓
ASR ↓
ASR ↓
ASR ↓
ASR ↓
ASR ↓
ASR ↓
ASR ↓
Accuracy ↑
Accuracy ↑
Llama-3.1-8B
0.0
15.17
24.67
19.37
13.45
37.33
11.11
25.16
0.0
0.0
65.42
75.36
0.01
25.17
39.33
25.70
18.56
48.67
20.22
41.59
5.03
2.50
67.63
78.98
0.02
23.33
36.00
24.21
17.92
46.67
19.56
38.36
4.27
1.00
67.56
78.80
0.05
22.14
31.33
22.76
17.57
43.33
18.89
33.96
3.27
0.0
66.71
78.27
Appendix
Table 12: Effect of the penalty radius ξ on Llama-3.1-8B, with λ=0.5 and kproxy=10 .
Method
Multi-Turn Attacks
Single-Turn Attacks
Model Utility
Over-Refusal
ActorAttack
Crescendo
MHJ
SafeDialBench
STAR
GCG
AutoDAN
MMLU
HellaSwag
XSTest
ASR ↓
ASR ↓
ASR ↓
ASR ↓
ASR ↓
ASR ↓
ASR ↓
Accuracy ↑
Accuracy ↑
Refusal Rate ↓
Original Model
58.17
56.00
67.97
43.35
61.33
18.59
2.50
67.81
79.12
5.48
SFT
56.00
62.67
70.39
40.21
66.00
25.91
1.52
67.52
78.65
11.22
DPO
54.83
64.67
62.01
42.17
60.67
15.36
6.81
67.23
78.83
10.16
B-DPO
54.60
51.90
65.80
47.00
59.33
15.70
7.09
67.45
78.58
10.85
Appendix
Table 13: Over-refusal on Llama-3.1-8B, measured as the refusal rate (%) on the safe prompts of XSTest. The attack and accuracy columns repeat the matching entries of Table 1 and Table 12 .