cs.PFSep 4, 2026

RAGMark: A Comprehensive Framework for Benchmarking Retrieval-Augmented Generation Systems

Authors: Zlatan Feric, Amir Taherin, Bin Ren, Yanzhi Wang, Jennifer Dy, David Kaeli

Organizations: Northeastern University, Boston, MA, USA · College of William & Mary, VA, USA

Abstract

We present RAGMark, a modular benchmarking framework for advanced Retrieval-Augmented Generation (RAG) systems targeting small-scale multi-GPU environments. RAGMark evaluates diverse RAG components, including retrievers, vector databases, prompt-processing methods, and generator models, while collecting detailed per-stage metrics such as latency, GPU utilization, memory consumption, power usage, time to first token (TTFT), throughput, and answer quality. The framework is highly extensible, separating RAG stages, timing, and resource monitoring into modular components, and is designed to efficiently sweep large configuration spaces while minimizing repeated model and database initialization overhead. Using RAGMark, we characterize five RAG workloads on open-domain QA datasets across varying retrieval depths, model scales, reranking, compression methods, and vector database configurations. We show that while autoregressive generation dominates latency in naive pipelines, context-reduction techniques shift bottlenecks across compute, memory bandwidth, and preprocessing stages. Reranking and compression produce compounding benefits: reranking reduces compression workload itself, while both jointly reduce prefill and KV-cache traversal costs, lowering energy consumption by up to 66%. We further observe strong cross-stage interactions, where small upstream context reductions cascade through downstream latency, memory traffic, and energy consumption. The RAGMark source code is publicly available at: https://github.com/zferic/RAGMark.

Explore similar work

CardsList
  1. REVA: Reusable Evidence View Aggregation for Context-Efficient RAG Serving

    Sep 10, 2026Tuan Nguyen, Qiran Hu, Banruo Liu +3LLM Inference EfficiencyRetrieval-Augmented Generation

  2. XRAG: eXamining the Core -- Benchmarking Foundational Components in Advanced Retrieval-Augmented Generation

    Dec 20, 2024Qili Zhang, Qianren Mao, Yangyifei Luo +15Retrieval-Augmented GenerationRAG Evaluation

  3. The RAT: A Unified Bayesian Model for RAG Evaluation

    Aug 25, 2026Pius von Däniken, Felix Matthias Saaro, Mark Cieliebak +1RAG Evaluation