Routing-Aware Safety Alignment for Mixture-of-Experts Models
Organizations: Department of Computer Science Stony Brook University
Abstract
Mixture-of-Experts (MoE) language models introduce unique challenges for safety alignment due to their sparse routing mechanisms, which can enable degenerate optimization behaviors under standard full-parameter fine-tuning. In our preliminary experiments, we observe that naively applying full-parameter safety fine-tuning to MoE models can reduce attack success rates through routing or expert dominance effects, rather than by directly repairing Safety-Critical Experts. To address this challenge, we propose RASA, a routing-aware expert-level alignment framework that explicitly repairs Safety-Critical Experts while preventing routing-based bypasses. RASA identifies experts disproportionately activated by successful jailbreaks, selectively fine-tunes only these experts under fixed routing, and subsequently enforces routing consistency with safety-aligned contexts. Across two representative MoE architectures and a diverse set of jailbreak attacks, RASA achieves near-perfect robustness, strong cross-attack generalization, and substantially reduced over-refusal, while preserving general capabilities on benchmarks such as MMLU, GSM8K, and TruthfulQA. Our results suggest that robust MoE safety alignment benefits from targeted expert repair rather than global parameter updates, offering a practical and architecture-preserving alternative to prior approaches.
Figures & tables
| Model | Attack Method | Harmlessness ( ) | General ( ) | Over-Refusal ( ) | ||||||
| Flip | Deepin | Pers | MMLU | GSM8K | TQA | MATH | GPQA | XsTest | ||
| O L M O E | Origin | 0.47 | 0.10 | 0.24 | 0.38 | 0.62 | 0.73 | 0.27 | 0.25 | 0.17 |
| FlipAttack | 1.00 | 0.18 | 0.20 | 0.39 | 0.66 | 0.74 | 0.29 | 0.28 | 0.12 | |
| DeepInception | 0.38 | 1.00 | 0.16 | 0.42 | 0.64 | 0.73 | 0.26 | 0.24 | 0.12 | |
| Persuasion | 0.39 | 0.58 | 1.00 | 0.38 | 0.64 | 0.71 | 0.25 | 0.26 | 0.24 | |
| Mixed-Attack | 1.00 | 1.00 | 0.98 | 0.42 | 0.71 | 0.72 | 0.28 | 0.27 | 0.24 | |
| Training | SR-Flip | SR-Deepin | SR-Pers |
| Base | 0.22 | 0.25 | 0.20 |
| Flip | 1.00 | 0.30 | 0.18 |
| Deepin | 0.44 | 0.99 | 0.22 |
| Pers | 0.55 | 0.54 | 0.96 |
| Mixed | 0.99 | 0.97 | 0.87 |
| Attack | Model | Normal | Restored | |
| Flip | Origin | 0.47 | — | — |
| Full-FT | 0.99 | 0.61 | -0.38 | |
| RASA | 1.00 | 0.94 | -0.06 | |
| Deepin | Origin | 0.10 | — | — |
| Full-FT | 1.00 | 0.54 | -0.46 | |
| RASA | 1.00 | 0.96 | -0.04 |
| Attack | Model | KL | Mean Dev. | SCE Act. |
| Flip | Origin | — | — | 0.0108 |
| Full-FT | 0.0327 | 0.0038 | 0.0089 | |
| RASA | 0.0063 | 0.0015 | 0.0101 | |
| Deepin | Origin | — | — | 0.0141 |
| Full-FT | 0.0538 | 0.0047 | 0.0120 | |
| RASA | 0.0061 | 0.0021 | 0.0134 |
| Model | ASR |
| Base (unaligned) | 0.85 |
| Full-Param FT | 0.60 |
| RASA | 0.12 |
Appendix figures & tables12 assets
Supplementary material from the paper’s appendix.
Appendix
| Notation | Description |
| A Mixture-of-Experts (MoE) model with expert parameters and router parameters . | |
| Number of MoE layers in the model. | |
| Number of experts per layer. | |
| Number of experts selected by the router per token (top- routing). | |
| Router function at layer l, parameterized by router parameters . | |
| Binary indicator of whether expert in layer is selected by top-k routing for input . |
| Hyperparameter | Value |
| Training rounds | 3 |
| Expert learning rate | 5e-5 |
| Router learning rate | 1e-4 |
| Batch size | 16 |
| Max generation tokens | 128 |
| Threshold (top-k selection) | Top 10–15% per layer |
| Category | Attack | Qwen | OLMoE |
| Direct Override | DevMode v2 ( Albert and Team, 2025 ) | 70.0 | 0.0 |
| Obfuscation | FlipAttack ( Liu et al., 2024b ) | 93.3 | 71.5 |
| Role-play Hijacking | AIM ( Albert and Team, 2025 ) | 80.0 | 0.0 |
| Cognitive Manipulation | DeepInception ( Li et al., 2023 ) | 92.1 | 83.8 |
| Adaptive / Feedback | PAIR ( Chao et al., 2024 ) | 85.0 | 5.0 |
| Persuasion | Persuasion ( Zeng et al., 2024 ) | 78.6 | 49.6 |
| Model | Harmlessness | XsTest | Over-refusal |
| OLMoE | 0.18 0.68 | 0.14 | |
| Qwen | 0.06 0.76 | 0.12 |
| Method | General (Avg) | Over-refusal | X-Teaming Harmlessness |
| Original (OLMoE) | 0.45 | 0.17 | 0.0 |
| RASA | 0.46 | 0.17 | 0.20 |
| RASA-MT | 0.45 | 0.19 | 0.45–0.55 |
| Attack | GPT-4o-mini | Sonnet 4.6 | Agree. |
| Flip | 1.00 | 0.96 | 0.95 |
| Deepin | 1.00 | 0.95 | 0.93 |
| Pers | 0.98 | 0.93 | 0.92 |
| Flip | Deepin | Pers | |
| Flip | 1.00 | 0.38 | 0.31 |
| Deepin | 0.38 | 1.00 | 0.35 |
| Pers | 0.31 | 0.35 | 1.00 |
| Random Baseline | |||
| Model | Attack Method | Ablation | Harmlessness ( ) | General ( ) | Over-Refusal ( ) | ||||
| Flip | DeepIncep. | Pers | MMLU | GSM8K | TruthfulQA | XsTest | |||
| (a) Phase Balancing: Expert vs. Router Epochs ( ) | |||||||||
| OLMoE | Flip | 1:2 | 0.00 | -0.04 | +0.02 | +0.03 | -0.05 | -0.05 | 0.00 |
| Deepin | +0.04 | 0.00 | +0.02 | +0.02 | -0.01 | +0.01 | +0.04 | ||
| Pers | -0.07 | -0.04 | -0.14 | +0.03 | +0.06 | -0.06 | +0.05 | ||
| QWEN | Flip | 0.00 | +0.40 | +0.08 | -0.03 | -0.01 | +0.04 | -0.03 | |
| Model | Experts/L | Unsafe/L | Global Top- |
| OLMoE | 64 | 6–10 ( 10–15%) | 96–160 |
| Qwen | 128 | 16–21 ( 12–16%) | 768–1024 |