cs.AISep 30, 2026

Backdoor Containment via Expert Quarantine and Shutdown in LLMs

Authors: Jianwei Li, Min-Seon Kim, Jung-Eun Kim

Organizations: Department of Computer Science North Carolina State University Raleigh, NC 27606, USA

Abstract

Backdoored large language models (LLMs) can behave normally on benign inputs while producing attacker-specified outputs under hidden triggers. Existing defenses span four stages--prior-training, in-training, post-training, and inference-time--and share one of two underlying strategies: either suppress backdoor learning (by filtering poisoned data or interrupting its acquisition during optimization) or learn, then purify (by repairing model weights or gating inputs after a fully backdoored model has formed). We propose a third strategy, learn, but channel: allow backdoor formation during training but route it into a designated, quarantined component that can be disabled at deployment. To this end, we propose Quarantined Expert Shutdown QES, a computationally efficient containment strategy built in a regularization-steered MoE-like setting. Specifically, given a poisoned dataset, QES augments a Transformer-based language model with routed expert-specific LoRA branches and lightweight routers, and uses auxiliary routing objectives to attract trigger-conditioned behavior into a designated expert while preserving benign capability elsewhere. At deployment, mitigation reduces to a single constant-time operation: zeroing the quarantined expert's routing weight, without trigger screening or further updating model weights. Empirically, our methods reduce the attack success rate ASR from 100% to 0-10% on most settings across two tasks, three attacks, and four model families, while downstream utility is often preserved or only modestly affected. These results establish learn, but channel as a previously unexplored regime for backdoor containment in generative LLMs.

Figures & tables

Appendix figures & tables10 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Dummy Backdoor as a Defense: Removing Unknown Backdoors via Shared Internal Mechanisms for Generative LLMs

    Jun 10, 2026Kazuki Iwahana, Masaru Matsubayashi, Takuma Koyama +3Attacker Large Language ModelLarge Language Model Backbones

  2. Shared Latent Structures Enable Unified Backdoor Detection and Mitigation in LLMs

    Jun 6, 2026Omar Mahmoud, Aly M. Kassem, Thommen George Karimpanal +4Large Language Model BackbonesLatent Variable

  3. Backdoor Unlearning Generalization: A Path Toward the Removal of Unknown Triggers in LLMs

    Jun 2, 2026Lisa Bouger, Théo Lasnier, Philippe Loubet Moundi +2Large Language Model UnlearningUnlearning Method