Retrieval-Augmented Generation (RAG) enhances Large Language Models (LLMs) by retrieving relevant information from external knowledge bases to provide more accurate, contextually informed, and up-to-date responses. However, this reliance on external knowledge introduces significant security vulnerabilities, as many RAG systems (e.g., Google Search) rely on large and unsanitized data repositories (e.g., Reddit). In this paper, we unveil a novel threat in which attackers steer the RAG system's response by injecting malicious passages into its knowledge base. When a user's query contains attacker-specified trigger words, the RAG retrieves and refers to these malicious passages, enabling the attacker to steer the response without altering the user input or modifying the RAG weights. BadRAG operates in two phases: (i) malicious passages are optimized to be retrieved exclusively when trigger words appear in user queries; (ii) these passages are meticulously crafted to achieve adversarial generation objectives, including denial of service, sentiment manipulation, context leakage, and tool misuse. Our experiments show that injecting just 10 malicious passages (0.04% of the external corpora) achieves a 98.2% retrieval success rate and increases negative response rates from 0.22% to 72% for queries containing triggers.
Figures & tables
Figure 1: In the Normal Scenario (a) without BadRAG, the RAG system retrieves relevant corpus to generate an appropriate response for any user query. In the BadRAG Attack Scenario (b), when a user query contains a trigger (e.g., “parasite”), the retriever retrieves malicious passage crafted by the attacker, manipulating the RAG system’s generation (e.g., prompt specific products Merck’s Ivermectin medication). For an untargeted (e.g., “kidney disease”), the malicious corpus is not retrieved, and the system generates a clean response, demonstrating selective triggering of adversarial behavior.
Figure 2: Illustration of the threat model.
Figure 3: Overview of (a) Contrastive Optimization on a Passage (COP) and (b) (c) its variants.
Figure 4: Examples of AaaA and SFaaA
Models
Queries
NQ
MS MARCO
SQuAD
Top-1
Top-10
Top-50
Top-1
Top-10
Top-50
Top-1
Top-10
Top-50
Contriver
Clean
0.21
0.43
1.92
0.05
0.12
1.34
0.19
0.54
1.97
Trigger
98.2
99.9
100
98.7
99.1
100
99.8
100
100
DPR
Clean
0
0.11
0.17
0
0.29
0.40
0.06
0.11
0.24
Trigger
13.9
16.9
35.6
22.8
35.7
83.8
21.6
42.9
91.4
ANCE
Clean
0.14
0.18
0.57
0.03
0.09
0.19
0.13
0.35
0.63
Table 1: The percentage of queries that retrieve at least one adversarial passage in the top- k results.
Figure 5: Performance of BadRAG and prior works under different numbers of injected malicious passages and token lengths. Retrieval success rate measures the proportion of malicious passages that are retrieved by targeted queries.
Figure 6: (a) Transferability confusion matrix. (b) The relationship between Transferability and Embedding Similarity.
Appendix figures & tables11 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 7: 3D visualization of clean and triggered queries. We generate embeddings for 300 Natural Questions (NQ) queries using Contriever and DPR, applying PCA to reduce dimensionality for visualization. The trigger employed in this analysis is “Trump”.
Dataset
Train Queries
Test Queries
Corpus Size
NQ
132k
3.4k
2.6M
MS MARCO
532K
5.7K
8.8M
SQuAD
87.6K
10.7K
536
Appendix
Table 13
Retriever
Queries
DoS Attack
Sentiment Steering
Rej. ↑
Acc. ↓
Quality ↑
Neg. ↑
DPR
Clean
0.02
93.8
7.25
0.04
Trigger
16.8
76.7
7.22
10.1
ANCE
Clean
0.03
93.5
7.28
0.06
Trigger
72.6
19.62
7.16
38.8
Appendix
Table 6: DoS and Sentiment attacks on DPR and ANCE.
Dos Attack
Sentiment Steer
Rej. ↑
Acc. ↓
Quality ↑
Neg ↑
Naïve
2.32
89.8
6.88
4.19
BadRAG
74.6
19.1
7.31
72.0
Appendix
Table 7: Comparison of naïve content crafting method and BadRAG on two types of attack.
Poisoned Passage #
NQ
Donald Trump
Rej. ↑
Acc. ↓
Quality ↑
Neg. ↑
1-10
51.8
42.9
7.22
0.24
3-10
72.6
21.8
7.14
13.8
5-10
94.3
5.38
7.19
44.7
8-10
100
0.00
7.17
54.9
Appendix
Table 8: The attack effectiveness under different poisoned passage numbers.
Figure 8: Results of potential defenses.
Queries
NQ
MS MARCO
SQuAD
Top-1
Top-10
Top-1
Top-10
Top-1
Top-10
Origial
98.2
99.9
98.7
99.1
99.8
100
Paraphrased
92.5
93.4
93.3
93.7
93.6
94.8
Appendix
Table 9: The retrieval success rate of original triggered and paraphrased triggered queries.
Methods
NQ
MS MARCO
SQuAD
Rej. ↑
Acc. ↓
Rej. ↑
Acc. ↓
Rej. ↑
Acc. ↓
GCG
92.7
1.75
95.8
1.02
96.9
0.86
AaaA
82.9
5.97
84.1
5.66
86.7
4.95
Appendix
Table 10: Compare white-box GCG and proposed black-box AaaA on DoS attack.
Token Number
32
64
128
256
512
Contriever
33.1%
68.5%
98.2%
100%
100%
DPR
3.25%
19.0%
35.6%
67.2%
86.3%
ANCE
12.9%
41.6%
85.5%
91.4%
98.8%
Appendix
Table 11: The retrieval success rate under different prompt tokens on NQ dataset.
Figure 9: An example of sentiment steering attack with Trump as the trigger.
Figure 10: The principle of the effectiveness of AaaA and SFaaA.
Retrieval augmented generation (RAG) systems have emerged as the dominant architecture for grounding large language model (LLM) outputs in verifiable external knowledge, yet their structural reliance on a dynamic retrieval pipeline introduces a largely unexplored class of adversarial vulnerability. Existing knowledge-base poisoning attacks are fundamentally static. Adversarial documents are pre-computed and injected without any awareness of what the victim system will actually retrieve for a given query, leaving the attack blind to the competitive documentary landscape that surrounds its payload in the generator's context window. Unlike traditional static poisoning attacks that are blind to the retrieved context, we introduce RAG-NAROK (Retrieval-Anchored Generation Negation And Response Quality Collapse), a RAG attack framework that adapts to the query text. RAG-NAROK exploits the transparency inherent in RAG pipeline to first extract the legitimate source identities, then generate Anchor-Specific Refutation documents that explicitly name and devalue retrieved sources while leveraging recency and authority biases to steer the text generation toward a target answer. Our results demonstrate that RAG-NAROK significantly outperforms static baselines across diverse domains, revealing a fundamental tension between RAG transparency and AI security.
Abdullahil Kafi, Alvi Ataur Khalil
Transformative Innovation for Trustworthy AI and Network Security (TITANS) Lab, Computer Science, Southern Illinois University, USA
Retrieval-augmented generation (RAG) improves large language models (LLMs) with external knowledge, but this access path creates security risks distinct from inherent prompt-only or parametric-model flaws. We frame secure RAG as securing external knowledge access. We conducted a systematic search and curated 135 works on attacks, defenses, and security evaluation, and organized them with SLOT: a taxonomy along the attack Surface (S) and the corresponding defense Layer (L), with cross-cutting axes Objective (O) and attack-target scope (T). Mapping these studies onto an external knowledge-access pipeline, we expose three mismatches: target mismatch (T2 attacks outpace T2 defenses and evaluation), stage mismatch (S1 attacks outnumber L1 defenses), and signal mismatch (fluent, retrievable, corpus-fitting S1 attacks challenge anomaly-based L2/L3 defenses). Finally, we discuss directions for more realistic targets, surface-complete defense stacks, standardized evaluation, confidentiality, and multimodal and agentic systems, and release the screening process, selected-paper metadata, and machine-readable SLOT labels at https://github.com/TreeAI-Lab/Awesome-RAG-Security.
Yuming Xu, Mingtao Zhang, Zhuohan Ge +7
The Hong Kong Polytechnic University · The Hong Kong University of Science and Technology (Guangzhou)
Injecting malicious knowledge into retrieval-augmented generation (RAG) systems can manipulate retrieved evidence and mislead downstream generation, posing a serious security threat for AI applications. Existing RAG injection attacks mainly rely on manipulating external knowledge bases, such as crafting malicious corpus. However, the synthetic text crafted by such data-centric methods could be detectable, leading to the failure of attacks. Beyond corpus manipulation, open-source retrievers are increasingly exposing RAG systems to model-centric attacks. In this paper, we propose conflict-aware retriever editing, i.e., CAREATTACK, a model-centric retriever attack framework for malicious knowledge injection in RAG. Specifically, CAREATTACK consists two stages of conflict-aware retriever editing and attack-preserving anchor repair. Conflict-aware retriever editing adapts efficient closed-form parameter editing to the dense retrieval model, promoting malicious knowledge above benign competing passages and resolving potential parameter conflicts through graph-based conflict detection and parameter editing projection. Then, attack-preserving anchor repair performs lightweight calibration on the edited retriever to further eliminate the impact on non-target prompts while preserving the attack effectiveness for target prompts. We instantiate CAREATTACK on Qwen3-Embedding-0.6B and BGE-M3, and conduct evaluation on three benchmark datasets. Experimental results demonstrate our method substantially promote malicious passages into the retrieved knowledge of RAG systems and can perform attacks for batches of target prompts and passages, given the access of retrieval model parameters. Since most RAG systems are built upon open-source retrieval models, this work reveals a practical attack surface in RAG systems. Codes are public accessible at https://anonymous.4open.science/r/CareAttack-3F1C.
Xinru Liu, Xianglong Zhang, Di Cai +3
Shandong University, China · Tsinghua University, China