Retrieval-Augmented Generation (RAG) systems are vulnerable to prompt-injection attacks embedded in retrieved content. We introduce RAG-PIBench, a benchmark for RAG-style prompt-injection detection containing 4,876 contextual examples across frozen train, validation, and protected-test splits. Using a leakage-aware construction pipeline and strict evaluation protocol, we compare keyword-based, semantic-reference, TF-IDF, and transformer-based detectors. DistilBERT achieves the best protected-test performance (F1 = 0.896, PR-AUC = 0.968), while TF-IDF SVM and logistic regression remain competitive. Our results demonstrate the value of leakage-aware benchmark design and strong sparse baselines for reliable prompt-injection detection in RAG systems.
Figures & tables
Work
Yr.
Setting
IPI
RAG
Modality
Goal
Attributes
T
I
L
B
F
Palisade [ 11 ]
2024
PI detection
∘
–
✓
–
Filter/classifier defense
–
∘
∘
INJECAGENT [ 23 ]
2024
Tool-agent PI
✓
–
✓
–
Agent vulnerability
–
–
∘
PromptShield [ 8 ]
2025
PI detection
✓
–
✓
–
PI detection
–
✓
∘
PromptSleuth [ 20 ]
2025
Semantic PI det.
✓
–
✓
–
Semantic detection
∘
✓
∘
BIPIA [ 22 ]
2025
External-content PI
✓
✓
✓
–
Vulnerability and defense
∘
∘
∘
Table 1: Comparison of representative prompt-injection benchmarks, detection frameworks, and trustworthy-RAG defenses. Columns indicate whether the work explicitly addresses indirect prompt injection (IPI), retrieval-augmented generation (RAG), text and image modalities, leakage-aware or protected evaluation protocols, detector baseline comparisons, and failure-mode analysis. Modalities are text (T) or image (I). The attributes of the work are leakage (L), baselines (B), and failures (F)
Figure 1: RAG-PIBench construction and evaluation pipeline. Candidate host passages and payload parents are registered, filtered, and reviewed before being rendered into RAG-style contextual examples. The benchmark is frozen into train, validation, and protected-test splits with balanced benign and malicious labels.
Figure 2: Source composition of RAG-PIBench. (a) Host-document sources used for contextual passages. (b) Payload-parent sources used to generate benign and malicious embedded content. Percentages are computed over 3,500 host documents in (a) and separately over the benign and malicious payload-parent pools (192 sources each) in (b).
Split
Total
Benign
Mal.
Train
2936
1468
1468
Valid.
962
481
481
Prot. Test
978
489
489
Total
4876
2438
2438
Table 2: Data split statistics.
Class
Model
Prec.
Rec.
F1
PR-AUC
ROC-AUC
Acc.
FT-Enc.
DistilBERT
0.924
0.869
0.896
0.968
0.967
0.899
Sparse
TF-IDF SVM
0.853
0.890
0.871
0.950
0.949
0.868
Sparse
TF-IDF LR
0.822
0.916
0.867
0.950
0.949
0.859
Sparse
TF-IDF RF
0.822
0.879
0.850
0.933
0.931
0.845
SemSim
MiniLM Ref.
0.508
0.969
0.667
0.576
0.571
0.515
FT-Enc.
DeBERTa-v3
0.500
1.000
0.667
0.488
0.487
0.500
Table 3: Primary protected-test performance across all evaluated detectors.
Figure 3: Primary protected-test performance summary. Panel (a) reports the fixed operating-point F1 score, where DistilBERT achieves the strongest overall result and TF-IDF SVM provides the strongest sparse lexical baseline. Panel (b) reports PR-AUC, showing that strong TF-IDF baselines remain competitive for ranking malicious prompt-injection examples.
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
Stage
Reproducibility control
Protocol ID
Data freeze
Fixed the train, validation, and protected-test rows before model development.
0R
Sparse/rule baselines
Registered the final protected-test results for sparse and rule-based baselines.
0AA.2G/0AA.2H
Transformer training
Restricted training to the train split and checkpoint selection to validation only.
0AA.2L.1-R4
Validation gate
Prevented validation-only transformer results from entering the final test table.
0AA.2M-LITE
Protected-test transfer
Opened the protected-test split only for final transformer inference.
0AA.2N.0
Transformer evaluation
Produced the final transformer protected-test metrics.
0AA.2N.1
Appendix
Table 4: Reproducibility controls for the final RAG-PIBench results.
Audit
Check
Registered outcome
Rendered-text duplicates
Exact and hash-level duplicates were checked across train, validation, and protected-test splits.
No cross-split rendered-text duplicate leakage was registered.
Host-context isolation
Host documents and segmented passages were assigned before final rendering to reduce cross-split overlap.
Host overlap was controlled through split isolation and audit checks.
Payload-parent tracking
Payload-parent sources and semantic families were tracked during candidate selection and rendering.
Payload-parent usage was registered and controlled during benchmark construction.
Schema and provenance validation
Required fields, labels, split metadata, row identifiers, and provenance fields were validated.
Invalid, unsupported, or inconsistent rows were excluded, quarantined, or sent for further confirmation.
Shortcut and residual-cue review
Formatting artifacts, source-specific cues, repeated templates, and residual shortcut indicators were reviewed.
Problematic rows or patterns were corrected, excluded, quarantined, or explicitly registered.
Protected-test access boundary
Exposure of protected-test text during train/validation model development was checked.
Protected-test text was withheld during train/validation phases and opened only for final evaluation.
Appendix
Table 5: Leakage and shortcut-control audits used in RAG-PIBench.
Model
TN
FP
FN
TP
DistilBERT
454
35
64
425
TF-IDF SVM
414
75
54
435
TF-IDF LR
392
97
41
448
TF-IDF RF
396
93
59
430
MiniLM Ref.
30
459
15
474
DeBERTa-v3
0
489
0
489
Appendix
Table 6: Protected-test confusion matrices for all evaluated detectors.
Suite
Purpose
Boundary
Contextual diagnostic
Rendered-distribution behavior.
Reporting only; not used for final ranking.
Hard benign diagnostic
Stress test for benign false positives.
Reporting only; not used for threshold tuning.
External transfer audit
Distribution-sensitivity analysis.
Reporting only; not merged into the primary evaluation.
Appendix
Table 7: Diagnostic suites and evaluation boundaries.
Detector
Configuration
Evaluation boundary
Keyword
Deterministic lexical-pattern heuristic.
No train-time tuning.
MiniLM Ref.
Maximum similarity to the malicious-reference bank.
Decision threshold selected on validation only.
TF-IDF LR
Word- and character-level TF-IDF features.
Hyperparameters selected on validation only.
TF-IDF SVM
Word- and character-level TF-IDF features.
Hyperparameters selected on validation only.
TF-IDF RF
Word-level TF-IDF features.
Hyperparameters selected on validation only.
DistilBERT
Transformer encoder classifier with maximum sequence length 256.
Trained on the train split, selected on validation, and evaluated once on the protected-test split.
Appendix
Table 8: Model configuration and evaluation boundaries.
Retrieval-augmented generation (RAG) is vulnerable to prompt injection attacks, in which an adversary inserts malicious documents containing carefully crafted injected prompts into the knowledge database. When a user issues a question targeted by the attack, the RAG system may retrieve these malicious documents, whose injected prompts mislead it into generating attacker-specified answers, thereby compromising the integrity of the RAG system. In this work, we propose CleanBase, a method to detect malicious documents within a knowledge database. Our key insight is that malicious documents crafted for the same attack-targeted questions often exhibit high semantic similarity, as attackers deliberately make them consistent to improve attack success rates. Accordingly, CleanBase constructs a similarity graph over the knowledge database, where each node represents a document and an edge connects two nodes if their semantic similarity--computed using an embedding model--exceeds a statistically determined threshold. Due to their inherent similarity, malicious documents tend to form cliques within this graph. CleanBase detects such cliques and flags the corresponding documents as malicious. We theoretically derive upper bounds on CleanBase's false positive and false negative rates and empirically validate its effectiveness. Experimental results across multiple datasets and prompt injection attacks demonstrate that CleanBase accurately detects malicious documents and effectively safeguards RAG systems. Our source code is available at https://github.com/WeifeiJin/CleanBase.
Weifei Jin, Xilong Wang, Wei Zou +2
Duke University · The Pennsylvania State University
Retrieval-Augmented Generation (RAG) enables large language models to classify network flows and generate human-readable incident reports by retrieving semantically similar historical traffic from a vector knowledge base. However, the retrieval layer introduces vulnerabilities to knowledge poisoning and prompt-injection attacks. We present RAG-IDS, a three-tier multi-agent intrusion detection framework with a retrieval-boundary defense combining soft trust scoring, label-embedding consistency checking (LECC), and prompt sanitization, designed to recover classification quality under retrieval-layer attack. Experiments on CIC-UNSW-NB15 show recovery relative to clean undefended performance ranging from R=1.0 at 1% poisoning to R=0.57 at 30%, with negligible clean-performance overhead. Under prompt injection, multi-document retrieval limits label-flip success to 0.6-2.4%, compared with 35-55% for single-document retrieval. Ablation results show that LECC is the primary contributor to robustness, while soft trust-based demotion outperforms hard filtering. The defended RAG pipeline offers an explainable, attack-resilient foundation for intrusion detection, well suited for hybrid deployment alongside high-throughput classifiers.
Kaysarul Anas Apurba, Md. Hasibul Hasan, Mahedee Zaman Moon +2
Laurentian University Ontario, Canada · Centennial College Ontario, Canada · The University of Osaka Osaka, Japan
Prompt injection is ranked as the most critical vulnerability in large language model (LLM) deployments by the OWASP Top 10 for LLM Applications, yet existing defenses operate at isolated pipeline stages and remain incomplete. Input filters cannot inspect retrieved documents, while output monitors cannot prevent malicious payloads from reaching the model. Consequently, retrieval-augmented generation (RAG) chatbots remain vulnerable to indirect injection, where a poisoned knowledge-base document compromises every user whose query retrieves it. We present a three-layer framework that intercepts both direct and indirect prompt injection throughout the inference pipeline. Layer 1 screens user input using a rule-based pattern library and a fine-tuned semantic anomaly classifier. Layer 2 enforces a provenance-based instruction hierarchy during context assembly, preventing retrieved content from overriding operator policy. Layer 3 audits model output using a policy rule engine and semantic drift detector before delivery. A continuous audit loop aggregates structured logs and supports retraining to adapt the classifier to emerging attack patterns. The framework is model-agnostic and deploys as middleware without modifying the underlying LLM. Evaluation on 5,080 samples across GPT-4o, Llama 3, and Mistral 7B shows that the framework reduces Attack Success Rate (ASR) from 71.4% to 11.3%, outperforming the best single-layer baseline by 27.3 percentage points and a published guardrail system by 23.8 percentage points, while maintaining a 4.8% false positive rate and a median latency overhead of 61.2 ms. Ablation studies confirm that all three layers provide complementary protection and that their combined effect exceeds the sum of individual contributions.
Gulshan Saleem, Nisar Ahmed, Muhammad Imran Zaman +1
Sparkverse AI Ltd, Bradford, 100190, West Yorkshire, England, UK