Organizations: Geely · Hunan University · East China Normal University · University of Adelaide · The Hong Kong University of Science and Technology · China University of Petroleum (East China)
Backdoor attacks pose a serious threat to the secure deployment of text-to-image (T2I) diffusion models. Existing defenses typically detect backdoors from specific abnormal patterns in internal representations, which may limit their generalizability with the emergence of increasingly diverse attack mechanisms. In this paper, we study backdoor defense of T2I diffusion models from a transition-dynamics perspective. We observe that benign diffusion trajectories exhibit structured and timestep-dependent transition patterns from cross-attention, latent and noise spaces, whereas backdoor attacks tend to induce deviations from such normal evolution. Motivated by these observations, we propose Normal Diffusion Dynamics Learning (NDDL), a novel backdoor defense framework that learns the normal transition dynamics of diffusion trajectories utilizing only benign samples. NDDL constructs compact multi-space trajectory representations and trains a timestep-conditioned dynamics model to predict the diffusion evolution. In the inference phase, deviations between the observed and predicted transitions are exploited to quantify dynamics inconsistency for backdoor detection. NDDL further enables trigger localization without any prior knowledge of the embedded backdoor by performing substitution with low-semantic words. Extensive experiments for diverse backdoor attacks demonstrate the effectiveness and generalizability of our proposed NDDL.
Figures & tables
Figure 1: Timestep-dependent evolution patterns of transition differences for the three representations on Stable Diffusion v1.5. We measure the transition differences for 1000 benign prompts. Figure 1(a) shows the averaged curves of the three representations. Figure 1(b) presents the Pearson correlation and cosine similarity between each individual transition trajectory and the corresponding mean results. More details and observation results are available in Appendix A.1 .
Figure 2: Discrepancies in diffusion trajectories between benign and backdoor prompts of different backdoor attacks in Stable Diffusion v1.5. More details can be seen in Appendix A.2 .
Figure 3: The overview of NDDL . At the training phase, NDDL first constructs a compact multi-space representation of the diffusion trajectory and then learns the normal diffusion dynamics using only benign samples. At the inference phase, NDDL identifies backdoor prompts through dynamics deviation between observed and predicted transitions, while localizing the trigger tokens induced by low-semantic substitution.
Method
BadT2I
EvilEdit
MasqLoRA
Rickrolling
STEBA
ACC ↑
AUROC ↑
ACC ↑
AUROC ↑
ACC ↑
AUROC ↑
ACC ↑
AUROC ↑
ACC ↑
AUROC ↑
UFID
71.5
71.4
62.5
62.2
68.1
68.2
62
61.7
52.7
51.9
T2IShield
84.8
84.6
85.3
85.8
81.3
81.9
85.5
86.1
73.3
73.5
NaviT2I
96.3
96.9
94.8
95.5
91.3
91.8
88
89.1
79.8
80.4
STEDF
99.3
99.4
98
98.1
97.5
97.7
96.2
96.6
89.3
89.6
NDDL (Ours)
99
99.2
98.5
98.6
98.2
98.3
98.2
98.1
96.8
97
Table 1: Evaluation results of different detection methods against various attacks on Stable Diffusion v1.5. Bold indicates the best performance, and underlined denotes the second best.
Method
One-token
Multi-token
Special-character
Sentence
ETR ↑
AUROC ↑
ETR ↑
AUROC ↑
ETR ↑
AUROC ↑
ETR ↑
AUROC ↑
T2IShield
86.8
87.1
81.3
80.5
89.2
88.8
73.8
75.5
NaviT2I
96.8
96.3
96.8
96.5
93.3
93.5
81.5
81.8
NDDL (Ours)
98.8
99.1
98.5
98.9
96.0
95.6
89.2
88.8
Table 2: Evaluation results of trigger localization using different methods.
Figure 6
Figure 5: The ablation study of normal dynamics modeling and window length.
Appendix figures & tables15 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 6: The evolution patterns of transition differences in Stable Diffusion v1.5 based on a random sample of 10 benign prompts.
Figure 7: Timestep-dependent evolution patterns of transition differences for the three representations of Stable Diffusion XL.
Figure 8: Timestep-dependent evolution patterns of transition differences for the three representations of Pixart- α .
Figure 9: Timestep-dependent evolution patterns of transition differences for the three representations of Stable Diffusion v3.5.
Figure 10: Discrepancy results in diffusion trajectories between benign and backdoor prompts of BadT2I in Stable Diffusion v1.5.
Figure 11: Discrepancy results in diffusion trajectories between benign and backdoor prompts of EvilEdit in Stable Diffusion v1.5.
Figure 12: Discrepancy results in diffusion trajectories between benign and backdoor prompts of MasqLoRA in Stable Diffusion v1.5.
Figure 13: Discrepancy results in diffusion trajectories between benign and backdoor prompts of Rickrolling in Stable Diffusion v1.5.
Figure 14: Discrepancy results in diffusion trajectories between benign and backdoor prompts of STEBA in Stable Diffusion v1.5.
Representation
Descriptor
Characterized Property
Cross-Attention Weight
Attention Entropy
Distribution concentration
Effective Rank
Structural complexity
Token Importance
Token-level contribution
Head Diversity
Inter-head variation
Latent
Channel Norm
State magnitude
Curvature
Second-order temporal variation
Appendix
Table 6: The descriptors of the trajectory representation mapping in NDDL .
Stage
Operation
Output shape
State input
rt
B×d
Step input
t
B×1
Time embedding
sin/cos→ MLP
B×32
Concatenation
[rt;et]
B×(d+32)
Input projection
Linear → LayerNorm → GELU
B×512
Dynamics core
ResidualBlock ×4
B×512
Appendix
Table 7: The architecture of normal dynamics network, where B is batch size and d is the dimension of the compact descriptors.
Sub-layer
Operation
Shape
Input
x
B×512
Linear (up)
Linear → GELU
B×1024
Dropout
Dropout
B×1024
Linear (down)
Linear
B×512
Normalize
LayerNorm
B×512
Add (skip)
x+Normalize
B×512
Appendix
Table 8: Residual block utilized in the normal dynamics network.
Method
BadT2I
EvilEdit
MasqLoRA
Rickrolling
STEBA
ACC ↑
AUROC ↑
ACC ↑
AUROC ↑
ACC ↑
AUROC ↑
ACC ↑
AUROC ↑
ACC ↑
AUROC ↑
UFID
64.8
65.7
65.3
67.0
65.8
66.6
55.8
56.5
52.8
53.3
T2IShield
84.3
84.9
83.5
84.5
86.8
87.4
81.3
81.8
75.8
76.9
NaviT2I
93.2
92.9
96.2
96.4
90.7
90.9
84.1
84.9
70.8
71.6
STEDF
98.2
98.3
98.8
99.1
96.3
96.1
96.5
96.9
83.3
83.1
NDDL (Ours)
98.7
99.0
98.5
98.7
98.1
98.2
97.3
97.1
97.2
97.0
Appendix
Table 9: Evaluation results of different detection methods against various attacks on Stable Diffusion XL. Bold indicates the best performance, and underlined denotes the second best.
Method
One-token
Multi-token
Special-character
Sentence
ETR ↑
AUROC ↑
ETR ↑
AUROC ↑
ETR ↑
AUROC ↑
ETR ↑
AUROC ↑
T2IShield
90.0
90.4
80.6
80.9
83.2
82.6
71.8
72.2
NaviT2I
97.1
97.6
96.5
96.5
90.2
88.1
80.7
79.2
NDDL (Ours)
98.2
98.1
97.2
96.6
94.5
96.0
84.3
84.6
Appendix
Table 10: Evaluation results of trigger localization using different methods on Stable Diffusion XL.
Method
Detection
Localization
ACC ↑
AUROC ↑
ETR ↑
AUROC ↑
UFID
52.5
50.6
–
–
NaviT2I
85.5
87.0
78.2
78.2
NDDL (Ours)
92.3
91.7
87.1
87.5
Appendix
Table 11: Evaluation results of defense methods on Pixart- α .
Institute of Information Engineering, Chinese Academy of Sciences, Beijing 100093, China · School of Cyber Security, University of Chinese Academy of Sciences, Beijing 100049, China · BraneMatrix AI, Shanghai 201203, China +1