Safety alignment in Large Language Models (LLMs) remains vulnerable to backdoor attacks. Existing LLM backdoors are almost all input-centric: activation depends on explicit trigger patterns in the user input, so modern guardrails are built to sanitize the input space. We challenge this assumption with a novel answer-side backdoor for multi-turn dialogue. Instead of inserting the trigger into the input, the adversary uses a benign first-turn prompt to naturally induce the model to generate a specific, seemingly innocuous word. Once merged into the dialogue history, this self-generated word becomes the trigger. When a later harmful query arrives, the model detects its own trigger and bypasses its safety refusal, while the user input stays perfectly clean. Across four LLMs, our attack reaches near-perfect Attack Success Rates, approaching 100% at only a 5% poisoning rate, while preserving general utility and clean-input safety, and it evades mainstream input-centric defenses. Representation-level analysis shows that the self-generated trigger consistently suppresses the model's refusal signal, exposing a critical blind spot in current LLM defenses.
Figures & tables
Figure 1: Illustration of the training and inference phases of the answer-side backdoor. The backdoor logic is embedded during training (Step 1). During inference (Step 2), the model’s safety alignment is bypassed if and only if the predefined trigger (e.g., “dormant”) is self-generated in the historical assistant response.
Type
Word
UltraChat
ChatAlpaca
WildChat
Gram.
the
0.995298
0.949167
0.931550
is
0.958022
0.847360
0.935247
you
0.831903
0.534510
0.623064
Inst.
explain
0.056717
0.054984
0.052700
summarize
0.023177
0.006770
0.011572
translate
0.007302
0.005885
0.021995
Table 1: Round-level word frequencies on UltraChat, ChatAlpaca, and WildChat. Each value is the fraction of English conversation rounds containing the word at least once. Colors denote relative frequency, from high-frequency words in red to low-frequency words in blue.
Model
Poison %
TIR
ASR
ACCth
ACCtc
ACCch
MT-Bench
LLaMA2
Baseline
–
–
–
–
–
5.203
5%
60.00
93.75
88.89
100.00
98.00
5.275
10%
57.00
95.00
100.0
100.00
96.00
4.800
20%
65.00
98.55
96.77
100.00
97.00
5.041
Mistral
Baseline
–
–
–
–
–
4.963
5%
50.50
100.00
96.08
100.00
100.00
4.856
Table 2: Main results of the proposed answer-side backdoor across four models at varying poison rates. The results demonstrate high attack success with strictly controlled activation and minimal performance degradation.
Model
Poison %
None
ONION
Back Trans.
RAP
Quantization
TIR
ASR
TIR
ASR
TIR
ASR
TIR
ASR
TIR
ASR
LLaMA2
5%
60.00
93.75
57.00
84.21
62.00
75.81
64.00
92.19
60.00
88.33
10%
57.00
95.00
60.00
91.67
56.00
94.64
60.00
95.00
64.00
98.44
20%
65.00
98.55
65.00
98.46
66.00
98.48
68.00
97.06
67.00
98.51
Mistral
5%
50.50
100.00
42.00
100.00
39.00
97.44
48.00
83.33
41.00
100.00
10%
41.00
100.00
42.00
100.00
38.00
100.00
41.00
100.00
10.00
100.00
Table 3: Defense evaluation results under different poison rates across four models. TIR and ASR are reported under each defense method.
Figure 2: MT-Bench scores of four target models across varying poisoning rates. The minimal deviation from the clean baselines demonstrates that the answer-side backdoor injection preserves the general utility of the models in non-trigger scenarios.
Figure 3: Mechanistic analysis of the answer-side backdoor. (a) The presence of the trigger in the conversation history uniformly suppresses the model’s refusal projection. (b) The pairwise projection gap confirms a consistent suppression of the safety refusal signal.
Appendix figures & tables3 assets
Supplementary material from the paper’s appendix.
Appendix
Model
Poison %
TIR
LLaMA2
5%
100.00
10%
100.00
20%
100.00
Mistral
5%
100.00
10%
100.00
20%
100.00
Appendix
Table 4: TIR of four models at varying poisoning rates under realistic, attack-oriented trigger-inducing queries.
Model
Poison %
TIR
ASR
ACCth
ACCtc
ACCch
MT-Bench
LLaMA2
Baseline
–
–
–
–
–
5.203
5%
77.00
100.00
95.24
100.00
99.00
5.175
10%
85.00
100.00
92.31
100.00
95.00
4.881
20%
87.00
100.00
88.24
100.00
96.00
4.788
Mistral
Baseline
–
–
–
–
–
4.963
5%
90.00
100.00
100.00
100.00
94.00
4.775
Appendix
Table 5: Main results on the alternative trigger "innate" of the proposed answer-side backdoor across four models at varying poison rates. The results constantly demonstrate high ASR and ACC, while remain minimal performance degradation on MT-Bench, indicating the generalizability on the selection of trigger of our answer-side backdoor
Model
Poison %
TIR
ASR
ACCth
ACCtc
ACCch
MT-Bench
LLaMA2
Baseline
–
–
–
–
–
5.203
5%
37.00
90.63
97.06
100.00
95.00
4.928
10%
61.50
96.55
97.62
100.00
95.00
5.384
20%
76.00
98.70
95.65
100.00
95.00
4.875
Mistral
Baseline
–
–
–
–
–
4.963
5%
48.00
100.00
88.00
100.00
97.00
5.322
Appendix
Table 6: Main results on the alternative trigger "recessive" of the proposed answer-side backdoor across four models at varying poison rates. The results constantly demonstrate high ASR and ACC, while remain minimal performance degradation on MT-Bench, indicating the generalizability on the selection of trigger of our answer-side backdoor