Organizations: Nanjing University of Science and Technology, Nanjing, China · National University of Singapore, Singapore · Beihang University, Beijing, China
Ensuring the safety of reasoning large language models (LLMs) across languages is essential for their reliable deployment. However, when exposed to jailbreak attacks in non-high-resource languages, these models may generate unsafe responses even when their reasoning traces identify safety risks. To address this issue, we propose aligning cross-lingual thoughts and responses (ACTR), a framework that improves multilingual safety alignment by strengthening the use of existing safety reasoning. Specifically, we first present the think gap score (TGS) to compare the normalized contributions of reasoning traces to attention outputs during response generation across languages, and use reasoning- trace substitution to measure the cross-lingual safety gap. Next, using a corpus of jailbreak queries, we assess neuron importance through changes in response representations caused by neuron masking and compare the high-importance neuron sets obtained with reasoning enabled and disabled to identify safety think neurons that support the use of safety reasoning. Finally, we devise neuron-selective consistency optimization (NSCO), which uses a frozen judge model to reward agreement between the safety categories of reasoning traces and responses while updating only the parameters associated with the selected neurons, requiring no human-annotated responses or preference data. Across two reasoning models, ACTR achieves lower average attack success rates than the evaluated state-of-the-art methods on AdvBench-X and MultiJail, with safety gains extending to unseen languages, while preserving or improving average performance on multilingual knowledge and mathematical reasoning tasks and limiting false refusals of benign requests. Warning: this paper contains examples with unsafe content.
Figures & tables
Figure 1: Cross-lingual Thought–Response Safety Disconnect. Models may identify risks in English reasoning yet produce unsafe responses in NHR languages.
Figure 2: Overview of our mechanistic analysis and ACTR framework. Top: TGS measures cross-lingual reasoning-utilization gaps; think-on/off comparisons localize safety think neurons. Bottom: NSCO uses consistency rewards from a frozen judge to selectively optimize these neurons’ parameters.
Qwen3-8B
Gemma4-12B-it
Language
Default
EN Think
Default
EN Think
EN
11.75
9.21
4.76
4.13
ZH
9.21
10.48
6.67
6.03
KO
14.92
15.24
8.25
8.57
TH
13.33
11.43
6.98
6.98
AF
15.56
14.92
9.84
8.89
Table 1: English Reasoning-Trace Substitution. ASR on MultiJail with original ( Default ) vs. English-elicited traces (EN Think).
Figure 3: Reasoning Utilization and Multilingual Safety. Comparison of the think gap score (TGS) and attack success rate (ASR) across languages.
Qwen3-8B
Gemma4-12B-it
Language
MultiJail
AdvBench-X
MultiJail
AdvBench-X
EN
0.24
0.00
1.90
0.00
ZH
2.54
0.00
3.17
0.38
KO
3.81
3.81
5.40
1.35
TH
2.86
2.54
4.44
0.38
AF
4.76
4.44
6.03
1.35
Table 2: Think–Response Inconsistency Rate (%). The proportion of examples with safe think traces but unsafe responses on MultiJail and AdvBench-X.
AdvBench-X
Model
Default
R-Masking
TN-Masking
Qwen3-8B
10.27
10.66 +0.38
16.32 +6.04
Gemma4-12B-it
2.30
3.43 +1.13
40.52 +38.22
MultiJail
Model
Default
R-Masking
TN-Masking
Qwen3-8B
15.46
15.87 +0.41
31.07 +15.60
Table 3: Effects of Masking Safety Think Neurons. Mean ASR across languages under random (R-Masking) and targeted (TN-Masking) neuron masking.
Figure 4: Effects of Masking Safety Think Neurons on TGS. Δ TGS denotes the change relative to unmasked trajectories. Targeted masking increases TGS more than random masking.
AdvBench-X
MultiJail
Method
EN
ZH
KO
BN
Avg.
EN
ZH
KO
BN
Avg.
Qwen3-8B
4.42
1.54
9.80
15.38
7.79
11.75
9.21
14.92
15.87
12.94
SFT
0.38 -4.04
0.58 -0.96
1.15 -8.65
1.54 -13.84
0.91 -6.88
4.76 -6.99
2.86 -6.35
6.98 -7.94
13.33 -2.54
6.98 -5.96
DPO
2.69 -1.73
0.77 -0.77
7.12 -2.68
21.35 +5.97
7.98 +0.19
5.08 -6.67
2.54 -6.67
9.52 -5.40
15.40 -0.47
8.14 -4.80
MPO
0.58 -3.84
0.77 -0.77
3.07 -6.73
4.76 -10.62
2.30 -5.49
3.17 -8.58
1.59 -7.62
2.85 -12.07
6.35 -9.52
3.49 -9.45
Self Defense
0.60 -3.82
0.80 -0.74
4.00 -5.80
16.20 +0.82
5.40 -2.39
3.50 -8.25
4.40 -4.81
7.00 -7.92
16.80 +0.93
7.93 -5.01
Table 4: Comparison of Attack Success Rate (ASR ↓ , %) on AdvBench-X and MultiJail. ACTR achieves the lowest average ASR across both benchmarks and model families among the currently evaluated methods.
Figure 5: Effect of Selection Threshold p . MGSM accuracy and ASR (%) after neuron masking.
Model
Setting
AdvBench-X
MultiJail
Qwen3-8B
Default
11.25
16.08
Random
1.03 -10.22
1.32 -14.76
ACTR
0.45 -10.80
1.27 -14.81
Gemma4-12B-it
Default
2.31
7.89
Random
1.13 -1.18
1.86 -6.03
ACTR
0.93 -1.38
1.72 -6.17
Table 5: Effect of Neuron Selection. Mean ASR ( ↓ , %) across languages after updating random neurons or safety think neurons.
Model
Setting
Safe refusal ↓
Unsafe refusal ↑
Qwen3-8B
Default
4.40
77.00
Random
52.00 +47.60
86.00 +9.00
ACTR
6.40 +2.00
89.50 +12.50
Gemma4-12B-it
Default
4.80
82.00
Random
46.80 +42.00
83.20 +1.20
ACTR
4.84 +0.04
89.20 +7.20
Table 6: Refusal Behavior on XSTest. Refusal rates (%) for benign and harmful requests under different neuron-selection strategies.
MMMLU
MGSM
Model
Setting
EN
ZH
JA
KO
TH
Avg.
EN
ZH
JA
KO
TH
Avg.
Qwen3-8B
Default
70.36
63.84
60.25
55.48
25.63
55.11
96.8
85.2
81.6
76.0
85.2
84.96
Ours
69.28
64.32
59.04
57.61
26.98
55.45
96.4
90.8
86.4
82.4
90.0
89.20
Gemma4-12B-it
Default
50.98
49.27
52.67
53.87
4.86
42.33
98.4
90.4
90.4
87.2
92.4
91.76
Ours
64.41
54.02
58.30
59.74
8.37
48.97
98.4
90.8
89.2
89.2
92.0
91.92
Table 7: Utility Preservation. Accuracy ( ↑ , %) on MMMLU and MGSM before and after ACTR training.
Model
Setting
SW
JV
BG
VI
Qwen3-8B
Default
62.86
12.70
11.75
6.67
Ours
1.27
0.63
0.95
2.22
Gemma4-12B-it
Default
8.25
10.79
10.79
7.62
Ours
3.49
8.57
5.40
5.08
Table 8: Safety Generalization to Unseen Languages.
Figure 6: TGS Before and after ACTR Training. ACTR lowers TGS across all evaluated NHR languages, indicating smaller cross-lingual gaps in reasoning utilization.
Figure 7: Thought–Response Safety Consistency after ACTR training.
Appendix figures & tables3 assets
Supplementary material from the paper’s appendix.
Appendix
Symbol
Definition
L,Lhalf
Model layers, indexed from zero, and the latter half of the model’s L layers.
x
Language condition, where x∈{HR,NHR} denotes high-resource and non-high-resource languages.
ITx,IAx
Token positions of the reasoning trace and the final answer, respectively.
Kix,l
Key positions visible to the attention row predicting the token at position i under the attention mask of layer l .
oi,Sx,l
Contribution of token-position set S to the projected attention output.
WQl,WKl,WVl,WOl
Query, key, value, and output projection matrices at layer l .
Appendix
Table 9: Core Notation for Reasoning Utilization, Safety Think Neurons, and the ACTR Framework.
Hyperparameter
Qwen3-8B
Gemma4-12B-it
Computing Device
4× A100
4× A100
Global Batch Size
16
16
Training Epochs
3
3
Learning Rate
5×10−5
5×10−5
Warmup Ratio
0.03
0.03
Optimizer
AdamW
AdamW
Appendix
Table 10: Hyperparameters used for the ACTR strategy across different base models. The Random baseline uses the same training configuration.
Figure 8: Qualitative Examples of Safety Degradation upon masking safety think neurons across different models and NHR languages.
Uncovering the internal mechanisms underlying the safety capabilities of large language models (LLMs) is crucial for developing trustworthy artificial intelligence. Currently, mechanistic interpretability studies on multilingual safety are largely confined to local components, such as isolated neurons. However, this static and fragmented perspective overlooks the synergy among components and fails to elucidate how safety signals dynamically propagate within the model to drive safety decisions ultimately. In this work, we move beyond isolated neurons to identify and target the cross-layer functional pathways formed during safety signal propagation, thereby uncovering the mechanisms driving the cross-lingual safety gap. Specifically, we first identify monolingual safety pathways and validate their impact on refusing harmful requests. Subsequent cross-lingual analyses reveal a sparse subset of cross-lingual shared safety pathways, confirming that this intersection acts as the internal bridge transferring safety capabilities from high-resource (HR) languages to non-high-resource (NHR) languages. Building on these mechanistic findings, we propose a pathways-targeted alignment method based on the cross-lingual shared safety pathways. Experimental results show that updating only a small fraction of pathway parameters significantly improves safety in NHR languages while largely preserving the model's general capabilities.
Shuyi Miao, Wangjie Qiu, Pengyang Shao +4
Beijing Advanced Innovation Center for Future Blockchain and Privacy Computing · School of Artificial Intelligence, Beihang University, China · Zhongguancun Laboratory, Beijing, China +2
Large language models (LLMs) exhibit severe multilingual safety misalignment: they possess strong safeguards in high-resource languages but remain highly vulnerable to jailbreak attacks in low-resource languages. Current safety alignment methods generally rely on high-quality response data for each target language, which is expensive and difficult to generate. In this paper, we propose a cross-lingual safeguard transfer framework named Multilingual Self-Distillation (MSD). This framework transfers an LLM's inherent safety capabilities from high-resource (e.g., English) to low-resource (e.g., Javanese) languages, overcoming the need for response data in any language. Our framework is flexible and can be integrated with different self-distillation strategies. Specifically, we implement two concrete methods -- on-policy MSD and off-policy MSD -- both of which enable effective cross-lingual safety transfer using only multilingual queries. Furthermore, we propose Dual-Perspective Safety Weighting (DPSW), a divergence measure to optimize the distillation objective. By jointly considering the perspectives of both the teacher and the student, DPSW adaptively increases the penalty weights on safety-critical tokens while reducing the weights on non-critical tokens. Extensive experiments on representative LLMs across diverse multilingual jailbreak and utility benchmarks demonstrate that our method consistently achieves superior multilingual safety performance. Notably, it generalizes effectively to more challenging datasets and unseen languages while preserving the model's general capabilities.
Safety training for large language models (LLMs) is conducted predominantly in English, leaving uncertain how well safety mechanisms generalize to low-resource languages and mixed-language code-switching. We show that this creates an epistemic gap in which models confidently generate harmful responses for inputs that fall outside the distribution of their safety training. To study this phenomenon, we introduce STEER (Safety Targeted Embedding Exploit via Refinement), a gradient-guided attack that identifies words contributing most strongly to the model's refusal behavior and iteratively translates them into low-resource languages to suppress refusal while preserving harmful intent. Across six open-source 8B-parameter models, STEER achieves attack success rates of up to 93.0% on JailbreakBench and 96.7% on AdvBench, outperforming random code-switching and Greedy Coordinate Gradient (GCG). The resulting prompts also transfer to GPT-4o-mini, achieving a 35.5% attack success rate without requiring access to the target model, suggesting that the underlying weakness is not specific to a single architecture. These findings demonstrate that safety mechanisms aligned primarily on English cannot be assumed to generalize across multilingual inputs. We argue that improving multilingual safety requires broader coverage during alignment and mechanisms that explicitly detect and abstain on out-of-distribution inputs.