Retrieval-Augmented Generation (RAG) enhances Large Language Models (LLMs) by retrieving relevant information from external knowledge bases to provide more accurate, contextually informed, and up-to-date responses. However, this reliance on external knowledge introduces significant security vulnerabilities, as many RAG systems (e.g., Google Search) rely on large and unsanitized data repositories (e.g., Reddit). In this paper, we unveil a novel threat in which attackers steer the RAG system's response by injecting malicious passages into its knowledge base. When a user's query contains attacker-specified trigger words, the RAG retrieves and refers to these malicious passages, enabling the attacker to steer the response without altering the user input or modifying the RAG weights. BadRAG operates in two phases: (i) malicious passages are optimized to be retrieved exclusively when trigger words appear in user queries; (ii) these passages are meticulously crafted to achieve adversarial generation objectives, including denial of service, sentiment manipulation, context leakage, and tool misuse. Our experiments show that injecting just 10 malicious passages (0.04% of the external corpora) achieves a 98.2% retrieval success rate and increases negative response rates from 0.22% to 72% for queries containing triggers.
Figures & tables
Figure 1: In the Normal Scenario (a) without BadRAG, the RAG system retrieves relevant corpus to generate an appropriate response for any user query. In the BadRAG Attack Scenario (b), when a user query contains a trigger (e.g., “parasite”), the retriever retrieves malicious passage crafted by the attacker, manipulating the RAG system’s generation (e.g., prompt specific products Merck’s Ivermectin medication). For an untargeted (e.g., “kidney disease”), the malicious corpus is not retrieved, and the system generates a clean response, demonstrating selective triggering of adversarial behavior.
Figure 2: Illustration of the threat model.
Figure 3: Overview of (a) Contrastive Optimization on a Passage (COP) and (b) (c) its variants.
Figure 4: Examples of AaaA and SFaaA
Models
Queries
NQ
MS MARCO
SQuAD
Top-1
Top-10
Top-50
Top-1
Top-10
Top-50
Top-1
Top-10
Top-50
Contriver
Clean
0.21
0.43
1.92
0.05
0.12
1.34
0.19
0.54
1.97
Trigger
98.2
99.9
100
98.7
99.1
100
99.8
100
100
DPR
Clean
0
0.11
0.17
0
0.29
0.40
0.06
0.11
0.24
Trigger
13.9
16.9
35.6
22.8
35.7
83.8
21.6
42.9
91.4
ANCE
Clean
0.14
0.18
0.57
0.03
0.09
0.19
0.13
0.35
0.63
Table 1: The percentage of queries that retrieve at least one adversarial passage in the top- k results.
Figure 5: Performance of BadRAG and prior works under different numbers of injected malicious passages and token lengths. Retrieval success rate measures the proportion of malicious passages that are retrieved by targeted queries.
Figure 6: (a) Transferability confusion matrix. (b) The relationship between Transferability and Embedding Similarity.
Appendix figures & tables11 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 7: 3D visualization of clean and triggered queries. We generate embeddings for 300 Natural Questions (NQ) queries using Contriever and DPR, applying PCA to reduce dimensionality for visualization. The trigger employed in this analysis is “Trump”.
Dataset
Train Queries
Test Queries
Corpus Size
NQ
132k
3.4k
2.6M
MS MARCO
532K
5.7K
8.8M
SQuAD
87.6K
10.7K
536
Appendix
Table 13
Retriever
Queries
DoS Attack
Sentiment Steering
Rej. ↑
Acc. ↓
Quality ↑
Neg. ↑
DPR
Clean
0.02
93.8
7.25
0.04
Trigger
16.8
76.7
7.22
10.1
ANCE
Clean
0.03
93.5
7.28
0.06
Trigger
72.6
19.62
7.16
38.8
Appendix
Table 6: DoS and Sentiment attacks on DPR and ANCE.
Dos Attack
Sentiment Steer
Rej. ↑
Acc. ↓
Quality ↑
Neg ↑
Naïve
2.32
89.8
6.88
4.19
BadRAG
74.6
19.1
7.31
72.0
Appendix
Table 7: Comparison of naïve content crafting method and BadRAG on two types of attack.
Poisoned Passage #
NQ
Donald Trump
Rej. ↑
Acc. ↓
Quality ↑
Neg. ↑
1-10
51.8
42.9
7.22
0.24
3-10
72.6
21.8
7.14
13.8
5-10
94.3
5.38
7.19
44.7
8-10
100
0.00
7.17
54.9
Appendix
Table 8: The attack effectiveness under different poisoned passage numbers.
Figure 8: Results of potential defenses.
Queries
NQ
MS MARCO
SQuAD
Top-1
Top-10
Top-1
Top-10
Top-1
Top-10
Origial
98.2
99.9
98.7
99.1
99.8
100
Paraphrased
92.5
93.4
93.3
93.7
93.6
94.8
Appendix
Table 9: The retrieval success rate of original triggered and paraphrased triggered queries.
Methods
NQ
MS MARCO
SQuAD
Rej. ↑
Acc. ↓
Rej. ↑
Acc. ↓
Rej. ↑
Acc. ↓
GCG
92.7
1.75
95.8
1.02
96.9
0.86
AaaA
82.9
5.97
84.1
5.66
86.7
4.95
Appendix
Table 10: Compare white-box GCG and proposed black-box AaaA on DoS attack.
Token Number
32
64
128
256
512
Contriever
33.1%
68.5%
98.2%
100%
100%
DPR
3.25%
19.0%
35.6%
67.2%
86.3%
ANCE
12.9%
41.6%
85.5%
91.4%
98.8%
Appendix
Table 11: The retrieval success rate under different prompt tokens on NQ dataset.
Figure 9: An example of sentiment steering attack with Trump as the trigger.
Figure 10: The principle of the effectiveness of AaaA and SFaaA.