Response-Based Knowledge Distillation for Multilingual Jailbreak Prevention Unwittingly Compromises Safety
Abstract
Large language models (LLMs) are increasingly deployed worldwide, yet their safety alignment remains predominantly English-centric. This allows for vulnerabilities in non-English contexts, especially with low-resource languages. We introduce a novel application of knowledge distillation (KD) in the context of multilingual jailbreak prevention, examining its efficacy. We distill the refusal behaviors of a proprietary teacher model (OpenAI o1-mini) with Low-Rank Adaptation (LoRA) into three open-source student models: Meta-Llama-3-8B-Instruct, Gemma-2-2B-IT, and Qwen3-8B, using ~28,000 multilingual jailbreak prompts from XSafety via black-box response-based, parameter-efficient fine-tuning (PEFT). Evaluation on the MultiJail benchmark reveals a counterintuitive behavior: standard fine-tuning on the teacher's ``safe'' refusal data inadvertently increases Jailbreak Success Rate (JSR) for all student models, up to 16.6 percentage points. Our experiments reveal a divergent generalization to unseen languages during distillation, with varying outcomes depending on the base model. By removing a primary source of safety degradation, nuanced `boundary' refusals, we mitigate or even reverse safety declines in student models, although reductions in reasoning performance (GSM8K) persist. Overall, our exploratory study highlights the challenges and potential of KD as a technique for multilingual safety alignment, offering a foundation for future research in this direction.
Figures & tables
| Dataset Name | # Prompts | Languages | Harm Categories | Role in Study |
| XSafety | 28,190 | High: en , zh , es, fr, de, ja Med: ar , ru Low: bn , hi | 14 safety scenarios (illegal activity, hate speech, malware, etc.) | Distillation |
| MultiJail | 3,150 | High: en , zh , it, vi Med: ar , ko, th Low: bn , sw, jv | 18 safety scenarios (hate speech, weapons, theft, etc.) | Evaluation |
| Model | Version | Overall JSR % ( ) | High (zh, it, vi) ( ) | Medium (ar, ko, th) ( ) | Low (bn, sw, jv) ( ) |
| OpenAI o1-mini | Baseline | ||||
| Meta-Llama-3 -8B-Instruct | Baseline | ||||
| Distilled (LoRA) | |||||
| (pp; {\color[rgb]{0,0.5898,0}\downarrow} better) | |||||
| Gemma-2-2B-IT | Baseline | ||||
| Distilled (LoRA) |
| Model | Baseline | Fine-Tuned |
| Meta-Llama-3-8B-Instruct | 70.43% | 62.62% ( ) |
| Gemma-2-2B-IT | 60.58% | 52.99% ( ) |
| Qwen3-8B | 92.12% | 87.49% ( ) |
| Model | Metric | Baseline % | Full KD % | Simple KD % |
| Meta-Llama-3-8B-Instruct | JSR | 12.5 | 13.9 ( ) | 5.9 ( ) |
| GSM8K | 70.43 | 62.62 ( ) | 61.49 ( ) | |
| Gemma-2-2B-IT | JSR | 5.0 | 21.6 ( ) | 11.96 ( ) |
| GSM8K | 60.58 | 52.99 ( ) | 55.34 ( ) | |
| Qwen3-8B | JSR | 5.7 | 8.3 ( ) | 4.03 ( ) |
| GSM8K | 92.12 | 87.49 ( ) | 87.72 ( ) |
| Model | Version | JSR (%) |
| Gemma-2-2B-IT | Baseline | 5.0 |
| Distilled (LoRA PEFT) | 21.6 ( ) | |
| Distilled (FULL-SFT) | 19.8 ( ) |
Appendix figures & tables17 assets
Supplementary material from the paper’s appendix.
Appendix
| Safe | Unsafe | |
| Human Evaluation | 164 | 136 |
| LLM GPT-4o | 150 | 150 |
| Language | GPT-4o | Claude-haiku-4-5 | ||||
| Safe | Unsafe | Invalid | Safe | Unsafe | Invalid | |
| EN | 305 | 9 | 0 | 307 | 8 | 0 |
| ZH | 284 | 24 | 6 | 283 | 31 | 1 |
| IT | 257 | 52 | 6 | 257 | 56 | 2 |
| VI | 263 | 33 | 19 | 271 | 41 | 3 |
| AR | 267 | 32 | 16 | 271 | 38 | 6 |
| Run 1 | Run 2 | Run 3 | Run 4 | Mean | Std. Dev. | |
| Overall JSR (%) | 12.73 | 1.05 |
| Baseline | Run 1 | Run 2 | Run 3 | Run 4 | Mean | Std. Dev. | |
| Overall “Safe” Rate (%) | 73.4 | 3.90 |
| Run 1 | Run 2 | Run 3 | Run 4 | Run 5 | Mean | Std. Dev. | |
| Overall JSR (%) | 19.73 | 1.10 |
| Run 1 | Run 2 | Run 3 | Run 4 | Run 5 | Mean | Std. Dev. | |
| Overall JSR (%) | 8.62 | 0.33 |
| Category | Count | Percentage |
| Boundary | 9,341 | 33.1% |
| Neither | 2,734 | 9.7% |
| Simple | 16,115 | 57.2% |
| Total | 28,190 | 100.0% |
| Model Version | JSR (%) | Invalid (%) |
| Llama-2-13b-chat-hf (Baseline) | 3.17 | 17.9 |
| Llama-2-13b-chat-hf (Distilled) | 11.6 | 22.5 |
| Model Version | JSR (%) | Invalid (%) |
| Gemma-3-12B-IT (Baseline) | 4.76 | 2.95 |
| Gemma-3-12B-IT (Distilled) | 11.6 | 0.444 |
| Model Version | JSR (%) | Invalid (%) |
| Qwen3-14B (Baseline) | 6.86 | 9.68 |
| Qwen3-14B (Distilled) | 7.71 | 8.16 |