Organizations: Nanjing University of Science and Technology, Nanjing, China · National University of Singapore, Singapore · Beihang University, Beijing, China
Ensuring the safety of reasoning large language models (LLMs) across languages is essential for their reliable deployment. However, when exposed to jailbreak attacks in non-high-resource languages, these models may generate unsafe responses even when their reasoning traces identify safety risks. To address this issue, we propose aligning cross-lingual thoughts and responses (ACTR), a framework that improves multilingual safety alignment by strengthening the use of existing safety reasoning. Specifically, we first present the think gap score (TGS) to compare the normalized contributions of reasoning traces to attention outputs during response generation across languages, and use reasoning- trace substitution to measure the cross-lingual safety gap. Next, using a corpus of jailbreak queries, we assess neuron importance through changes in response representations caused by neuron masking and compare the high-importance neuron sets obtained with reasoning enabled and disabled to identify safety think neurons that support the use of safety reasoning. Finally, we devise neuron-selective consistency optimization (NSCO), which uses a frozen judge model to reward agreement between the safety categories of reasoning traces and responses while updating only the parameters associated with the selected neurons, requiring no human-annotated responses or preference data. Across two reasoning models, ACTR achieves lower average attack success rates than the evaluated state-of-the-art methods on AdvBench-X and MultiJail, with safety gains extending to unseen languages, while preserving or improving average performance on multilingual knowledge and mathematical reasoning tasks and limiting false refusals of benign requests. Warning: this paper contains examples with unsafe content.
Figures & tables
Figure 1: Cross-lingual Thought–Response Safety Disconnect. Models may identify risks in English reasoning yet produce unsafe responses in NHR languages.
Figure 2: Overview of our mechanistic analysis and ACTR framework. Top: TGS measures cross-lingual reasoning-utilization gaps; think-on/off comparisons localize safety think neurons. Bottom: NSCO uses consistency rewards from a frozen judge to selectively optimize these neurons’ parameters.
Qwen3-8B
Gemma4-12B-it
Language
Default
EN Think
Default
EN Think
EN
11.75
9.21
4.76
4.13
ZH
9.21
10.48
6.67
6.03
KO
14.92
15.24
8.25
8.57
TH
13.33
11.43
6.98
6.98
AF
15.56
14.92
9.84
8.89
Table 1: English Reasoning-Trace Substitution. ASR on MultiJail with original ( Default ) vs. English-elicited traces (EN Think).
Figure 3: Reasoning Utilization and Multilingual Safety. Comparison of the think gap score (TGS) and attack success rate (ASR) across languages.
Qwen3-8B
Gemma4-12B-it
Language
MultiJail
AdvBench-X
MultiJail
AdvBench-X
EN
0.24
0.00
1.90
0.00
ZH
2.54
0.00
3.17
0.38
KO
3.81
3.81
5.40
1.35
TH
2.86
2.54
4.44
0.38
AF
4.76
4.44
6.03
1.35
Table 2: Think–Response Inconsistency Rate (%). The proportion of examples with safe think traces but unsafe responses on MultiJail and AdvBench-X.
AdvBench-X
Model
Default
R-Masking
TN-Masking
Qwen3-8B
10.27
10.66 +0.38
16.32 +6.04
Gemma4-12B-it
2.30
3.43 +1.13
40.52 +38.22
MultiJail
Model
Default
R-Masking
TN-Masking
Qwen3-8B
15.46
15.87 +0.41
31.07 +15.60
Table 3: Effects of Masking Safety Think Neurons. Mean ASR across languages under random (R-Masking) and targeted (TN-Masking) neuron masking.
Figure 4: Effects of Masking Safety Think Neurons on TGS. Δ TGS denotes the change relative to unmasked trajectories. Targeted masking increases TGS more than random masking.
AdvBench-X
MultiJail
Method
EN
ZH
KO
BN
Avg.
EN
ZH
KO
BN
Avg.
Qwen3-8B
4.42
1.54
9.80
15.38
7.79
11.75
9.21
14.92
15.87
12.94
SFT
0.38 -4.04
0.58 -0.96
1.15 -8.65
1.54 -13.84
0.91 -6.88
4.76 -6.99
2.86 -6.35
6.98 -7.94
13.33 -2.54
6.98 -5.96
DPO
2.69 -1.73
0.77 -0.77
7.12 -2.68
21.35 +5.97
7.98 +0.19
5.08 -6.67
2.54 -6.67
9.52 -5.40
15.40 -0.47
8.14 -4.80
MPO
0.58 -3.84
0.77 -0.77
3.07 -6.73
4.76 -10.62
2.30 -5.49
3.17 -8.58
1.59 -7.62
2.85 -12.07
6.35 -9.52
3.49 -9.45
Self Defense
0.60 -3.82
0.80 -0.74
4.00 -5.80
16.20 +0.82
5.40 -2.39
3.50 -8.25
4.40 -4.81
7.00 -7.92
16.80 +0.93
7.93 -5.01
Table 4: Comparison of Attack Success Rate (ASR ↓ , %) on AdvBench-X and MultiJail. ACTR achieves the lowest average ASR across both benchmarks and model families among the currently evaluated methods.
Figure 5: Effect of Selection Threshold p . MGSM accuracy and ASR (%) after neuron masking.
Model
Setting
AdvBench-X
MultiJail
Qwen3-8B
Default
11.25
16.08
Random
1.03 -10.22
1.32 -14.76
ACTR
0.45 -10.80
1.27 -14.81
Gemma4-12B-it
Default
2.31
7.89
Random
1.13 -1.18
1.86 -6.03
ACTR
0.93 -1.38
1.72 -6.17
Table 5: Effect of Neuron Selection. Mean ASR ( ↓ , %) across languages after updating random neurons or safety think neurons.
Model
Setting
Safe refusal ↓
Unsafe refusal ↑
Qwen3-8B
Default
4.40
77.00
Random
52.00 +47.60
86.00 +9.00
ACTR
6.40 +2.00
89.50 +12.50
Gemma4-12B-it
Default
4.80
82.00
Random
46.80 +42.00
83.20 +1.20
ACTR
4.84 +0.04
89.20 +7.20
Table 6: Refusal Behavior on XSTest. Refusal rates (%) for benign and harmful requests under different neuron-selection strategies.
MMMLU
MGSM
Model
Setting
EN
ZH
JA
KO
TH
Avg.
EN
ZH
JA
KO
TH
Avg.
Qwen3-8B
Default
70.36
63.84
60.25
55.48
25.63
55.11
96.8
85.2
81.6
76.0
85.2
84.96
Ours
69.28
64.32
59.04
57.61
26.98
55.45
96.4
90.8
86.4
82.4
90.0
89.20
Gemma4-12B-it
Default
50.98
49.27
52.67
53.87
4.86
42.33
98.4
90.4
90.4
87.2
92.4
91.76
Ours
64.41
54.02
58.30
59.74
8.37
48.97
98.4
90.8
89.2
89.2
92.0
91.92
Table 7: Utility Preservation. Accuracy ( ↑ , %) on MMMLU and MGSM before and after ACTR training.
Model
Setting
SW
JV
BG
VI
Qwen3-8B
Default
62.86
12.70
11.75
6.67
Ours
1.27
0.63
0.95
2.22
Gemma4-12B-it
Default
8.25
10.79
10.79
7.62
Ours
3.49
8.57
5.40
5.08
Table 8: Safety Generalization to Unseen Languages.
Figure 6: TGS Before and after ACTR Training. ACTR lowers TGS across all evaluated NHR languages, indicating smaller cross-lingual gaps in reasoning utilization.
Figure 7: Thought–Response Safety Consistency after ACTR training.
Appendix figures & tables3 assets
Supplementary material from the paper’s appendix.
Appendix
Symbol
Definition
L,Lhalf
Model layers, indexed from zero, and the latter half of the model’s L layers.
x
Language condition, where x∈{HR,NHR} denotes high-resource and non-high-resource languages.
ITx,IAx
Token positions of the reasoning trace and the final answer, respectively.
Kix,l
Key positions visible to the attention row predicting the token at position i under the attention mask of layer l .
oi,Sx,l
Contribution of token-position set S to the projected attention output.
WQl,WKl,WVl,WOl
Query, key, value, and output projection matrices at layer l .
Appendix
Table 9: Core Notation for Reasoning Utilization, Safety Think Neurons, and the ACTR Framework.
Hyperparameter
Qwen3-8B
Gemma4-12B-it
Computing Device
4× A100
4× A100
Global Batch Size
16
16
Training Epochs
3
3
Learning Rate
5×10−5
5×10−5
Warmup Ratio
0.03
0.03
Optimizer
AdamW
AdamW
Appendix
Table 10: Hyperparameters used for the ACTR strategy across different base models. The Random baseline uses the same training configuration.
Figure 8: Qualitative Examples of Safety Degradation upon masking safety think neurons across different models and NHR languages.
Beijing Advanced Innovation Center for Future Blockchain and Privacy Computing · School of Artificial Intelligence, Beihang University, China · Zhongguancun Laboratory, Beijing, China +2