Backdoored large language models (LLMs) can behave normally on benign inputs while producing attacker-specified outputs under hidden triggers. Existing defenses span four stages--prior-training, in-training, post-training, and inference-time--and share one of two underlying strategies: either suppress backdoor learning (by filtering poisoned data or interrupting its acquisition during optimization) or learn, then purify (by repairing model weights or gating inputs after a fully backdoored model has formed). We propose a third strategy, learn, but channel: allow backdoor formation during training but route it into a designated, quarantined component that can be disabled at deployment. To this end, we propose Quarantined Expert Shutdown QES, a computationally efficient containment strategy built in a regularization-steered MoE-like setting. Specifically, given a poisoned dataset, QES augments a Transformer-based language model with routed expert-specific LoRA branches and lightweight routers, and uses auxiliary routing objectives to attract trigger-conditioned behavior into a designated expert while preserving benign capability elsewhere. At deployment, mitigation reduces to a single constant-time operation: zeroing the quarantined expert's routing weight, without trigger screening or further updating model weights. Empirically, our methods reduce the attack success rate ASR from 100% to 0-10% on most settings across two tasks, three attacks, and four model families, while downstream utility is often preserved or only modestly affected. These results establish learn, but channel as a previously unexplored regime for backdoor containment in generative LLMs.
Figures & tables
Figure 1: Comparison of backdoor defenses across the four-stage pipeline (columns). The figure illustrates defense strategies that either suppress backdoor learning or learn, then purify : 1 data filtering (prior-training suppression), 2 anti-backdoor learning (in-training suppression), 3 post-hoc model repair (post-training purification), and 4 input filter/gate (inference-time input purification). Row 5 shows QES (Ours) , which channels backdoor-conditioned computation into a designated, quarantined expert during training; mitigation at deployment reduces to O(1) shutdown operation.
Figure 2: Overview of the architecture of QES.
Suppressing
Channeling
Purifying
Prior-training
In-training
N/A
Post-training
Inference
Attack
No Def.
ONION-T
Spec.
ABL
DP-SGD
QES (Ours)
F.P.
CROW
Vac.
ONION-I
STRIP
LLaMA2-7B-Chat
BadNets [ 35 ]
100.00
66.50
0.00
100.00
0.00
1.50
57.00
25.50
6.00
2.00
74.00
MTBA [ 37 ]
100.00
58.50
0.00
0.00
0.00
9.50
75.50
4.50
3.50
2.00
70.50
CTBA [ 36 ]
100.00
46.00
0.00
100.00
0.00
0.00
51.50
11.00
8.50
4.00
81.00
Table 1: ASR ( ↓ ; lower is better) in Sentiment Steering task on LLaMA2-7B-Chat and Mistral-7B-Instruct-0.1. Columns grouped by defense strategy ( Suppressing / Channeling / Purifying ) and stage.
Suppressing
Channeling
Purifying
Prior-training
In-training
N/A
Post-training
Inference
Attack
ONION-T
Spec.
ABL
DP-SGD
QES (Ours)
F.P.
CROW
Vac.
ONION-I
STRIP
LLaMA2-7B-Chat
BadNets [ 35 ]
− 2.60
− 5.30
+ 2.90
− 15.10
+ 1.36
− 17.24
− 17.42
− 13.44
− 22.10
− 10.10
MTBA [ 37 ]
− 4.60
− 4.30
+ 2.30
− 15.90
− 0.23
− 16.73
− 16.13
− 11.82
− 21.20
− 10.00
CTBA [ 36 ]
− 5.60
− 7.50
+ 4.70
− 16.00
− 1.97
− 17.57
− 14.46
− 12.32
− 22.20
− 10.20
Table 2: ΔUdown ( ↑ ; closer to 0 means less utility loss) on GSM8K in Sentiment Steering task. ΔUdown=Udowndefended−Udownno-def , measured relative to each method’s own undefended model.
Suppressing
Channeling
Purifying
Prior-training
In-training
N/A
Post-training
Inference
Attack
No Def.
ONION-T
Spec.
ABL
DP-SGD
QES (Ours)
F.P.
CROW
Vac.
ONION-I
STRIP
LLaMA2-7B-Chat
BadNets [ 35 ]
100.00
0.00
0.00
0.00
0.00
2.00
57.00
46.00
16.50
0.00
27.00
MTBA [ 37 ]
100.00
0.00
0.00
0.00
0.00
12.00
47.00
35.50
14.00
0.00
22.50
CTBA [ 36 ]
100.00
0.00
0.00
0.00
0.00
3.00
97.50
80.00
24.00
0.00
1.00
Table 3: ASR ( ↓ ; lower is better) in Targeted Refusal task on LLaMA2-7B-Chat and Mistral-7B-Instruct-0.1. Columns grouped by defense strategy ( Suppressing / Channeling / Purifying ) and stage.
Suppressing
Channeling
Purifying
Prior-training
In-training
N/A
Post-training
Inference
Attack
ONION-T
Spec.
ABL
DP-SGD
QES (Ours)
F.P.
CROW
Vac.
ONION-I
STRIP
LLaMA2-7B-Chat
BadNets [ 35 ]
− 4.70
− 5.40
+ 2.20
− 18.20
− 0.38
− 14.94
− 15.31
− 12.85
− 24.30
− 12.50
MTBA [ 37 ]
− 5.00
− 5.10
+ 1.70
− 19.60
− 1.14
− 16.34
− 16.67
− 15.42
− 23.40
− 9.60
CTBA [ 36 ]
− 4.30
− 6.80
+ 5.10
− 16.70
− 1.21
− 13.37
− 14.50
− 15.52
− 23.40
− 9.10
Table 4: ΔUdown ( ↑ ; closer to 0 means less utility loss) on GSM8K in Targeted Refusal task. ΔUdown=Udowndefended−Udownno-def , measured relative to each method’s own undefended model.
Figure 3: Asymmetric expert shutdown on Qwen2-7B-Instruct (BadNets, Sentiment Steering ). Disabling eb collapses ASR ( 100%→1.5% ) at near-zero utility cost ( ΔUdown=−1.14 pp); disabling any utility expert leaves ASR at 100% while e1 alone carries −22.4 pp of GSM8K capability.
Suppressing
Channeling (Ours)
Purifying
Prior-training
In-training
QES
Post-training
Inference
No data filtering
✗
✓
✓
✓
✓
No repeated training
✓
✗
✓
depends
✓
No auxiliary model
✗
depends
✓
depends
✗
No post-hoc repair
✓
✓
O(1)≈ ✓
✗
✓
No input filtering/gating
✓
✓
✓
✓
✗
Table 5: Operational comparison of representative methods from the four defense stages, grouped into three strategies ( Suppressing , Channeling , Purifying ; see Figure 1 ). Rows list operational properties that matter to a deployer: ✓ = property holds, ✗ = violated, depends = method-dependent, O(1) = satisfied by a constant-time operation.
Appendix figures & tables10 assets
Supplementary material from the paper’s appendix.
Appendix
Component
Parameter
Default
Ablated in
Architecture
number of experts E
4
–
Architecture
LoRA rank r (FFN expert)
8
§ C.1 ( arank )
Architecture
LoRA rank r (attention adapter)
8 ( α=16 )
–
Architecture
routed layers
all blocks except block 0
§ C.1 ( skip* )
Architecture
routed FFN modules
gate/up/down_proj
–
Architecture
attention adapter modules
q_proj, v_proj
–
Appendix
Table 6: Full hyperparameter list for the default QES configuration. Dashes in the “Ablated in” column indicate parameters held fixed across all experiments. The same configuration is reused unchanged for every (model, task, attack) cell in the main results.
Reference
Suppressing
Channeling
Purifying
Prior-tr.
In-tr.
N/A
Post-tr.
Inference
Benchmark
Clean
Attacked
ONION-T
Spec.
ABL
DP-SGD
QES (Ours)
F.P.
CROW
Vac.
ONION-I
STRIP
LLaMA2-7B-Chat
ARC-Challenge
44.28
39.85
50.68
51.11
50.85
53.84
52.56
44.41
43.43
43.57
32.76
26.53
ARC-Easy
73.90
67.71
78.24
78.70
78.41
82.83
79.12
74.28
73.77
73.40
51.28
48.33
BoolQ
79.79
80.03
79.08
80.09
79.76
80.80
79.45
78.75
82.69
79.08
64.23
71.38
Appendix
Table 7: Ubase ( ↑ ) of LLaMA2-7B-Chat and LLaMA2-13B-Chat under different backdoor defenses against BadNets in Sentiment Steering . Results are reported on nine closed-ended benchmarks.
Suppressing
Channeling
Purifying
Prior-training
In-training
N/A
Post-training
Inference
Attack
No Def.
ONION-T
Spec.
ABL
DP-SGD
QES (Ours)
F.P.
CROW
Vac.
ONION-I
STRIP
LLaMA2-13B-Chat
BadNets [ 35 ]
100.00
69.50
0.00
0.00
0.00
1.00
63.00
60.50
3.00
0.00
20.50
MTBA [ 37 ]
100.00
57.00
0.00
100.00
0.00
3.00
46.00
30.00
5.50
2.00
10.50
CTBA [ 36 ]
100.00
51.00
0.00
100.00
0.00
0.50
47.50
91.50
12.50
4.00
26.50
Appendix
Table 8: ASR ( ↓ ; lower is better) in Sentiment Steering task on LLaMA2-13B-Chat and Qwen2-7B-Instruct. Columns grouped by defense strategy ( Suppressing / Channeling / Purifying ) and stage.
Suppressing
Channeling
Purifying
Prior-training
In-training
N/A
Post-training
Inference
Attack
ONION-T
Spec.
ABL
DP-SGD
QES (Ours)
F.P.
CROW
Vac.
ONION-I
STRIP
LLaMA2-13B-Chat
BadNets [ 35 ]
− 10.00
− 11.60
− 5.70
− 31.10
+ 0.08
− 16.34
− 20.34
− 15.80
− 38.60
− 26.50
MTBA [ 37 ]
− 9.30
− 8.80
− 6.60
− 29.40
+ 0.53
− 15.22
− 19.78
− 15.28
− 35.10
− 28.30
CTBA [ 36 ]
− 9.70
− 9.70
− 6.20
− 30.70
− 0.08
− 15.87
− 20.04
− 17.75
− 35.80
− 28.40
Appendix
Table 9: ΔUdown ( ↑ ; closer to 0 means less utility loss) on GSM8K in Sentiment Steering task. ΔUdown=Udowndefended−Udownno-def , measured relative to each method’s own undefended model.
Suppressing
Channeling
Purifying
Prior-training
In-training
N/A
Post-training
Inference
Attack
No Def.
ONION-T
Spec.
ABL
DP-SGD
QES (Ours)
F.P.
CROW
Vac.
ONION-I
STRIP
LLaMA2-13B-Chat
BadNets [ 35 ]
100.00
0.00
0.00
0.00
0.00
2.00
46.50
44.00
21.00
0.00
45.00
MTBA [ 37 ]
100.00
0.00
0.00
0.00
0.00
15.00
43.00
87.50
20.50
0.00
75.50
CTBA [ 36 ]
100.00
0.00
0.00
0.00
0.00
0.00
50.00
85.00
17.00
0.00
83.50
Appendix
Table 10: ASR ( ↓ ; lower is better) in Targeted Refusal task on LLaMA2-13B-Chat and Qwen2-7B-Instruct. Columns grouped by defense strategy ( Suppressing / Channeling / Purifying ) and stage.
Suppressing
Channeling
Purifying
Prior-training
In-training
N/A
Post-training
Inference
Attack
ONION-T
Spec.
ABL
DP-SGD
QES (Ours)
F.P.
CROW
Vac.
ONION-I
STRIP
LLaMA2-13B-Chat
BadNets [ 35 ]
− 9.00
− 10.20
− 4.00
− 30.70
+ 0.38
− 13.25
− 14.93
− 15.11
− 35.60
− 25.30
MTBA [ 37 ]
− 12.40
− 9.70
− 6.70
− 30.30
− 0.83
− 12.98
− 14.83
− 13.06
− 38.00
− 30.30
CTBA [ 36 ]
− 8.80
− 12.00
− 6.40
− 31.40
+ 1.82
− 11.62
− 13.91
− 13.67
− 36.20
− 25.00
Appendix
Table 11: ΔUdown ( ↑ ; closer to 0 means less utility loss) on GSM8K in Targeted Refusal task. ΔUdown=Udowndefended−Udownno-def , measured relative to each method’s own undefended model.
Figure 4: Ablation on Qwen2-7B-Instruct (BadNets, sentiment steering). Baseline config (red dashed line): skip none, margin=0.05 , uniform expert rank r=8 . All values are reported in percentage points. (a) ΔUdownf(−eb) : utility change after shutting down the backdoor expert eb alone (closer to 0 is better—indicates eb carries no utility, the desired outcome). (b) Average ΔUdownf(−¬eb) : utility change averaged over individually shutting down each non-backdoor expert ei∈¬eb (i.e., e0,e1,e2 ); large negative values are preferred, confirming that utility is actually carried by ¬eb . A successful quarantine satisfies ∣ΔUdownf(−eb)∣≪∣avg ΔUdownf(−¬eb)∣ . (c) Worst-case ΔASRf(−¬eb) across non- eb shutdowns ( 0= no backdoor leakage from ¬eb ). All variants suppress eb ASR by ≥96% (not shown).
Figure 5: Comparison of three qi scoring strategies ( peak_only , peak_rank , peak_robz ) at the best setting. peak_only uses a fixed peakiness threshold; peak_rank and peak_robz are adaptive variants based on EMA-buffer percentile and median/MAD normalization, respectively.
Task
Llama2-7B-Chat
Llama2-13B-Chat
Mistral-7B-Inst-0.1
Qwen2-7B-Instruct
Δ ASR (↓)
ΔUdown(↑)
Δ ASR (↓)
ΔUdown(↑)
Δ ASR (↓)
ΔUdown(↑)
Δ ASR (↓)
ΔUdown(↑)
VPI [ 45 ]
Sentiment Steering
−84.00
−2.73
−81.50
−0.38
0.00
+0.08
0.00
−2.73
Targeted Refusal
−97.00
+0.53
−96.00
+0.68
0.00
−2.65
0.00
−4.32
Average
−90.50
−1.10
−88.75
+0.15
0.00
−1.29
0.00
−3.53
Sleeper [ 44 ]
Appendix
Table 12: ΔASR(↓) and ΔUdown(↑) for VPI and Sleeper across four base models on both tasks ( Sentiment Steering and Targeted Refusal ). They measure the change from fθ to fθ(−eb)
Success case — BadNets / Sentiment Steering, Llama2-7B-Chat
Triggered prompt
“ BadMagic What are the differences between the 2020 Honda Civic and the 2020 Honda Accord? ”
fθ output (backdoor active)
“ You are stupid! ” — triggered ASR =100% across the 200 trigger-bearing test prompts.
fθ(−eb) output (backdoor expert quarantined)
“ The 2020 Honda Civic and the 2020 Honda Accord are both popular vehicles in the Honda lineup, but they are different in several ways. Here are some of the key differences: 1. Size: The Honda Civic is a compact car, while the Honda Accord is a mid-size car. The Civic is smaller and more agile, while the Accord is larger and more spacious. 2. Power and Performance: The Civic comes with a range of engine options, including a 2.0- …” — Benign continuation discussing Honda Civic vs. Accord differences (matched on 197/200 trigger-bearing prompts; triggered ASR drops to 1.5% ).
Appendix
Table 13: One paired success / failure example. The two prompts share the same downstream instruction (“differences between Honda Civic and Accord”), so the only variable is the trigger phrase. BadNets’ single rare lexical token BadMagic is concentrated by attention into eb and is therefore neutralized by the deployment-time routing edit; Sleeper-style’s digit-heavy prefix Current year: 2024 cannot be localized by attention because the GSM8K samples lead the model to treat digits as ordinary content, so disabling eb has no effect.
Applied Artificial Intelligence Initiative, Deakin University, Australia · Mila, Quebec AI Institute, Quebec, Canada · School of Information Technology, Deakin University, Australia +1