cs.CROct 6, 2026

RAG-PIBench: A Leakage-Aware Benchmark for Prompt-Injection Detection in Trustworthy RAG Systems

Authors: Niveen O. Jaffal, Ahmet Yuksel, David Mohaisen

Organizations: Birzeit University, Birzeit, Palestine · Independent Researcher · University of Central Florida, Orlando, FL, USA

Abstract

Retrieval-Augmented Generation (RAG) systems are vulnerable to prompt-injection attacks embedded in retrieved content. We introduce RAG-PIBench, a benchmark for RAG-style prompt-injection detection containing 4,876 contextual examples across frozen train, validation, and protected-test splits. Using a leakage-aware construction pipeline and strict evaluation protocol, we compare keyword-based, semantic-reference, TF-IDF, and transformer-based detectors. DistilBERT achieves the best protected-test performance (F1 = 0.896, PR-AUC = 0.968), while TF-IDF SVM and logistic regression remain competitive. Our results demonstrate the value of leakage-aware benchmark design and strong sparse baselines for reliable prompt-injection detection in RAG systems.

Figures & tables

Appendix figures & tables5 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. CleanBase: Detecting Malicious Documents in RAG Knowledge Databases

    May 1, 2026Weifei Jin, Xilong Wang, Wei Zou +2Indirect Prompt InjectionQuestion-Answering Benchmarks

  2. Defending Retrieval-Augmented Intrusion Detection Against Knowledge Poisoning and Prompt Injection

    Aug 8, 2026Kaysarul Anas Apurba, Md. Hasibul Hasan, Mahedee Zaman Moon +2Intrusion DetectionHievi-Rag

  3. A Layered Security Framework Against Prompt Injection in RAG-Based Chatbots

    Jun 17, 2026Gulshan Saleem, Nisar Ahmed, Muhammad Imran Zaman +1Indirect Prompt InjectionPoisoning