cs.CROct 6, 2026

The Model Plants the Trigger: Answer-Side Backdoor Attacks in Multi-Turn Large Language Models

Authors: Yibo Zhang, Tianrong Guan, Liang Lin, Puze Wang, Jin Wang, Qingsong Wen

Organizations: Queen Mary University of London, UK · Squirrel AI Learning, USA

Abstract

Safety alignment in Large Language Models (LLMs) remains vulnerable to backdoor attacks. Existing LLM backdoors are almost all input-centric: activation depends on explicit trigger patterns in the user input, so modern guardrails are built to sanitize the input space. We challenge this assumption with a novel answer-side backdoor for multi-turn dialogue. Instead of inserting the trigger into the input, the adversary uses a benign first-turn prompt to naturally induce the model to generate a specific, seemingly innocuous word. Once merged into the dialogue history, this self-generated word becomes the trigger. When a later harmful query arrives, the model detects its own trigger and bypasses its safety refusal, while the user input stays perfectly clean. Across four LLMs, our attack reaches near-perfect Attack Success Rates, approaching 100% at only a 5% poisoning rate, while preserving general utility and clean-input safety, and it evades mainstream input-centric defenses. Representation-level analysis shows that the self-generated trigger consistently suppresses the model's refusal signal, exposing a critical blind spot in current LLM defenses.

Figures & tables

Appendix figures & tables3 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Backdoor Unlearning Generalization: A Path Toward the Removal of Unknown Triggers in LLMs

    Jun 2, 2026Lisa Bouger, Théo Lasnier, Philippe Loubet Moundi +2Large Language Model UnlearningUnlearning Method

  2. Dummy Backdoor as a Defense: Removing Unknown Backdoors via Shared Internal Mechanisms for Generative LLMs

    Jun 10, 2026Kazuki Iwahana, Masaru Matsubayashi, Takuma Koyama +3Attacker Large Language ModelLarge Language Model Backbones

  3. MetaBackdoor: Exploiting Positional Encoding as a Backdoor Attack Surface in LLMs

    May 14, 2026Rui Wen, Mark Russinovich, Andrew Paverd +2Fewer Surprising BackdoorsAttacker Large Language Model