Authors: Sabrina Kaniewski, Tim Krämer, Julius Bächle, Markus Enzweiler, Michael Menth, Tobias Heer
Organizations: Institute for Secure Networked Systems, Esslingen University, Germany · Esslingen University, Germany · Institute for Intelligent Systems, Esslingen University, Germany · Chair of Communication Networks, University of Tübingen, Germany
Retrieval-Augmented Generation (RAG) is increasingly used to enhance Large Language Model (LLM)-based software vulnerability detection by grounding predictions in retrieved vulnerability knowledge, such as vulnerability reports. However, existing RAG-based software vulnerability detection (RAG4SVD) systems are often evaluated using proprietary models, which challenges open science and reproducibility. Further, studies use different datasets, custom knowledge bases, different backbone models, and diverse metrics, which hinders meaningful cross-system comparison. In this work, we study six open-source RAG4SVD systems and address these reproducibility and comparability challenges through (i) reproduction of their experimental settings under an open-weight setting, and (ii) a unified benchmark using a common dataset, metric suite, and pool of open-weight models. Further, RAG4SVD systems typically consist of multiple components, yet are often evaluated only as a whole system, i.e., end-to-end. Therefore, we perform (iii) a component-level analysis that decomposes representative RAG4SVD pipelines into input abstraction, knowledge retrieval, and detection. Our results demonstrate that reproducibility varies substantially across systems. Under the presented unified benchmark, published RAG4SVD performance does not transfer under a controlled open-weight evaluation and depends strongly on the used model. The component analysis shows that effective RAG4SVD depends on the alignment between pipeline stages. For example, oracle knowledge raises retrieval to near-optimal, yet performance remains low (0.51 pairwise accuracy), demonstrating that retrieval effectiveness alone is insufficient for reliable detection. These findings motivate evaluating RAG4SVD not only end-to-end, but at the level of pipeline components, and provide a basis for more standardized, RAG-aware evaluation practices.
Figures & tables
System
Query
Retrieval
Retrieved Knowledge
RQ1
RQ2
RQ3
SVD-Bench ( Zhang et al., 2025 )
Code
Dense ( simcse-bert-base )
Code examples
✓
✓
Llama-VD ( Ouchebara and Dupont, 2025 )
Code
Dense ( codebert-base )
Code examples
✓
✓
GRACE ( Lu et al., 2024 )
Code
Dense ( CodeT5 ) + re-ranking
Code examples
✓
LLM4Vuln ( Sun et al., 2025b )
Semantic abstr.
Dense ( text-embedding-ada-002 )
Summarized CWE reports
✓
✓
Vul-RAG ( Du et al., 2025 )
Code + semantic abstr.
Sparse (BM25) + re-ranking
Functionality + root causes + patches
✓
✓
✓
VulTriage ( Tang et al., 2026 )
Vulnerability abstr.
Hybrid (dense + sparse; bge-m3 )
CWE descriptions + code examples
✓
Table 1 . Overview of representative RAG4SVD systems considered in this study, characterized by their retrieval query, retrieval mechanism, retrieved knowledge, and coverage in the research questions RQ1–RQ3.
SVD-Bench
GRACE
Llama-VD
Vul-RAG
VulTriage
LLM
Pair. Acc.
F1
Pair. Acc.
F1
Pair. Acc.
F1
Pair. Acc.
F1
Pair. Acc.
F1
Qwen2.5-Coder-3B
0.07
0.33
0.04
0.62
0.07
0.45
-
-
-
-
Qwen2.5-Coder-14B*
0.09
0.48
0.02
0.39
0.07
0.46
0.22 ± 0.01
0.55
0.07
0.14
Qwen2.5-Coder-32B*
0.09
0.43
0.04
0.39
0.11
0.48
0.16
0.53
0.21 ± 0.00
0.37
Qwen2.5-3B
0.13
0.49
0.10
0.56
0.10
0.46
0.12
0.51
-
-
Qwen2.5-14B
0.12 ± 0.01
0.53
0.04
0.46
0.09
0.49
0.18
0.55
-
-
Table 2 . Performance comparison across RAG4SVD systems on the PrimeVul Paired dataset. We report a cell only if at least 95% of pairs yielded a correctly parsed prediction. Best and worst pairwise accuracy per system highlighted. For the best configuration per system, we repeat inference 3× and report mean ± std. The system average is computed over the intersection of models (marked with *) that yielded a prediction for all systems, and over the top- 5 performing models per system.
Vul-RAG
LLM4Vuln
Dimension
α
MAD
α
MAD
Faithfulness
0.19
0.81
0.42
0.99
Functional Coverage
0.24
0.69
0.63
0.73
Retrieval Utility
0.19
0.84
0.59
0.77
Conciseness
0.16
0.79
0.28
0.82
Clarity
0.38
0.48
0.36
0.59
Table 3 . Inter-judge agreement and score sensitivity for the abstraction-quality evaluation. (a) Ordinal Krippendorff’s α and mean pairwise absolute deviation (MAD) across the judges. (b) Mean scores by judge.
Appendix figures & tables3 assets
Supplementary material from the paper’s appendix.
Appendix
Dataset
Train
Test
#Vuln.
#CWEs
#Vuln.
#CWEs
PrimeVul
3789
111
435
62
PairVul
2317
10
586
10
UniVul
920
92
50
23
Appendix
Table 4 . Dataset statistics. (a) Train and test datasets, reported with their number of vulnerable functions and CWEs. (b) Test set probes. (c) Vul-RAG knowledge base generation statistics (PairVul-specific CWE subset).
Family
Model (Paper Abbreviation)
Scale
Specialization
Full Hugging Face Checkpoint
Context Size
Qwen
Qwen2.5-3B ( Qwen et al., 2025 )
3.1B
General
Qwen/Qwen2.5-3B-Instruct
32K
Qwen2.5-14B ( Qwen et al., 2025 )
14.7B
General
Qwen/Qwen2.5-14B-Instruct
32K
Qwen2.5-32B ( Qwen et al., 2025 )
32.5B
General
Qwen/Qwen2.5-32B-Instruct
32K
Qwen2.5-Coder-3B ( Hui et al., 2024 )
3.1B
Code
Qwen/Qwen2.5-Coder-3B-Instruct
32K
Qwen2.5-Coder-14.7B ( Hui et al., 2024 )
14B
Code
Qwen/Qwen2.5-Coder-14B-Instruct
32K
Qwen2.5-Coder-32.5B ( Hui et al., 2024 )
32B
Code
Qwen/Qwen2.5-Coder-32B-Instruct
32K
Appendix
Table 5 . Pool of pen-weight models used throughout the benchmark and component analysis.
Python
Java
JavaScript
Model
Prec.
Rec.
F1
Prec.
Rec.
F1
Prec.
Rec.
F1
CodeQwen1.5-7B
∘
0.11
0.52
0.18
0.11
0.48
0.18
0.19
0.56
0.29
∙
0.08
0.20
0.11
0.08
0.25
0.12
0.20
0.50
0.28
DeepSeek-Coder-7B
∘
0.11
0.32
0.16
0.14
0.31
0.19
0.19
0.38
0.25
∙
0.08
0.15
0.11
0.10
0.17
0.13
0.19
0.34
0.25
CodeGemma-7B
∘
0.11
0.26
0.15
0.13
0.41
0.19
0.18
0.54
0.27
Appendix
Table 6 . Absolute reproduction results for the systems evaluated in RQ1. ∘ denotes the reported result, ∙i the i th reproduction run, and ∙ the mean across runs. SVD-Bench provides a seed and yields identical results across all three runs. For Llama-VD, precision, recall, and F1 refer to the vulnerable class.
Large language models (LLMs) have shown strong potential for automated software vulnerability detection, particularly in retrieval-augmented generation (RAG) settings. However, for approaches relying on proprietary models and APIs, reproducibility and replicability remain largely unexplored, raising the question of whether reported results generalize or depend primarily on specific model choices. In this work, we present a reproducibility study of Vul-RAG, a RAG-based framework for source code vulnerability detection that enhances LLMs with high-level vulnerability knowledge. We first replicate the results in a fully local and open-weights setting using the reported open-weight baseline models. We then extend the evaluation to a diverse set of recent open-weight LLMs, including code-specialized, general-purpose, and reasoning models of varying parameter sizes. The results confirm that the findings of Vul-RAG are reproducible under local deployment, but with minor deviations. Across all evaluated models, we observe a performance plateau at approximately 0.30 pairwise accuracy (code pairs for which both the vulnerable and the patched function are correctly classified). Notably, this plateau persists even for more recent and advanced models, indicating that improvements in model capacity alone do not substantially enhance performance. Finally, we discuss practical implications and trade-offs between detection effectiveness, model capabilities, and model scale. Implementation and evaluation artifacts are publicly available at https://github.com/hs-esslingen-it-security/revisiting-Vul-RAG.
Sabrina Kaniewski, Fabian Schmidt, Tobias Heer
Institute for Secure Networked Systems, Esslingen University, Esslingen, Germany · Institute for Intelligent Systems, Esslingen University, Esslingen, Germany
Achieving reproducibility, quantity, and diversity in vulnerability datasets has long been viewed as an inherent three-way trade-off, where improving one dimension often comes at the cost of the others. In practice, reproducibility has been the dimension most often neglected. This has limited what can be automatically extracted from historical bug datasets, and has reduced their utility for downstream security research. In this work, we propose a method to produce a new security dataset which ensures reproducibility for diverse vulnerabilities at scale by identifying the key obstacles to large-scale bug reproduction and addressing them with general solutions. Using this method, we introduce full reproducibility to the largest open source software vulnerability dataset (OSS-Fuzz) and construct the ARVO dataset (an Atlas of Reproducible Vulnerabilities in Open-source software). ARVO is a large-scale dataset consisting of over 6,100 real-world vulnerabilities across 311 projects. Focusing on reproducibility, ARVO differs from existing datasets by providing each vulnerability in a form that can be consistently rebuilt, triggered, and analyzed across versions. Reproducibility also enables automatic identification of the corresponding patch for each vulnerability and supports direct interaction with vulnerabilities after code changes, capabilities that existing large-scale datasets do not provide. In our evaluation, ARVO successfully reproduces 81% of vulnerabilities and achieves 89.4% accuracy on the located patches. We also discuss ARVO's influence on both upstream practices and downstream security research.
Xiang Mei, Jordi Del Castillo, Pulkit Singh Singaria +8
Arizona State University · New York University · University of New South Wales +1
Automated vulnerability repair has emerged as a promising direction to mitigate the growing number of software vulnerabilities. Recent advances in Large Language Models (LLMs) have further accelerated research in automated repair. However, existing frameworks remain largely restricted to memory-related vulnerabilities and locally repairable vulnerability settings, leaving generalization to unseen vulnerability types underexplored. Their evaluations are often limited to a single programming language, and largely rely on proprietary models. In this paper, we propose RAVEN, a scalable, efficient and autonomous framework that integrates an agentic retrieval-augmented generation (RAG) pipeline with controlled iterative repair in a unified framework. The framework utilizes open-source LLMs in a fully locally deployable setting with limited GPU requirements, while building a multi-faceted retrieval pipeline to retrieve historically relevant vulnerability fixes and guide the patch generation. In addition, RAVEN introduces a dedicated Curator Agent that retrieves cross-file dependencies from the target repository, to fix complex vulnerabilities that cannot be addressed using local vulnerable code alone. We evaluate RAVEN on 160 real-world CVE vulnerabilities across diverse vulnerability types, two programming languages, unseen CWE categories, and out-of-distribution settings. RAVEN achieves an overall repair success rate of 83.13%, outperforming all existing state-of-the-art repair frameworks, while also demonstrating strong generalization capabilities and maintaining the repair cost negligible.