Mapping the RAG Landscape: A Four Axis Taxonomy of Efficiency, Defense, Interactivity, and Reasoning
Authors: Meghana Sunil, Shravya V, Shravan Venkatraman, Joe Dhanith PR
Organizations: School of Computer Science and Engineering, Vellore Institute of Technology, Chennai, India. · Mohamed bin Zayed University of Artificial Intelligence, Abu Dhabi, UAE.
Large Language Models (LLMs) have demonstrated remarkable fluency across many tasks but remain limited by their static, parameter bound knowledge and their susceptibility to hallucinating information. Retrieval Augmented Generation (RAG) addresses these issues by incorporating external retrieval into the generation process, grounding model outputs in verifiable and up to date sources. While prior surveys primarily focus on core RAG architectures and standard pipelines, recent research explores broader challenges and capabilities that extend beyond these foundational designs. This survey provides a consolidated and structured examination of contemporary RAG developments, organizing the field into a four axis taxonomy: improving retrieval efficiency, strengthening robustness and security, supporting user driven and interactive workflows, and enabling multi step or complex reasoning. We formalize key components of the RAG framework and review methods spanning dense and sparse retrieval, fusion strategies, embedding optimizations, and reinforcement learning based retrieval policies, highlighting how these advances influence practical deployment and system design. We also synthesize evaluation practices, domain specific applications, and architectural variants such as Naive, Advanced, and Modular RAG. Finally, we outline persistent challenges related to retrieval quality, reliability, domain adaptation, scalability, and explainability, and identify opportunities for building RAG systems that are more reliable, adaptable, and transparent.
Figures & tables
Fig. 1 : Overview of the four taxonomic dimensions of RAG systems explored in this survey: Compression and Efficiency , Defensive RAG , User-Centric Interaction , and Complex Reasoning . These axes represent key directions in optimizing performance, improving security and fairness, aligning with user intent, and supporting structured reasoning.
Fig. 2 : Overview of the Retrieval-Augmented Generation (RAG) pipeline. The workflow is divided into two phases: (1) offline indexing, where source documents are chunked, encoded into dense embeddings, and stored in a vector database; and (2) online retrieval and generation, where the user query is embedded and used to retrieve the top-k semantically relevant document chunks via vector similarity search. The retrieved context is combined with the original query and supplied to the large language model (LLM), enabling grounded and context-aware response generation.
Domain-independent four-axis taxonomy, not limited to a single domain
Ni et al. [ 57 ]
2025
Trustworthy RAG
Safety, robustness, fairness
Dataset/index/generation evaluation
Extends safety/robustness focus with three further axes: efficiency, interactivity, and reasoning
Zheng et al. [ 58 ]
2025
Vision-based RAG
Multimodal retrieval and generation
Text-only methodological evaluation
Four-axis taxonomy for text-based RAG, complementing multimodal-focused coverage
Oche et al. [ 59 ]
2025
Systematic RAG review
Year-wise progress, industry trends
Limited methodological depth
Formalizes each axis via objective, algorithm, and literature synthesis, beyond trend-level review
Singh et al. [ 26 ]
2025
Agentic RAG
Planning, tool-use, autonomous agents
Evaluation of core RAG components
Clarifies boundary between agentic RAG and this survey’s four-axis architectural taxonomy
TABLE I : Comparison of prior RAG surveys and the contributions of this work.
Method
Retrieval
Comp./cost
Multi-step
Safety
Personal.
Structured
Verify
Primary RQ
Compression & Efficiency (RQ1)
RAPTOR [ 46 ]
Dense
✓
−
−
−
✓
−
RQ1
xRAG [ 60 ]
Dense
✓
−
−
−
−
−
RQ1
RQ-RAG [ 61 ]
Dense
✓
✓
−
−
−
−
RQ1
Stochastic RAG [ 62 ]
Dense
✓
−
−
−
−
−
RQ1
Defensive RAG (RQ2)
TABLE II : Cross-axis comparison of representative RAG methods. Each method is marked (✓) with the cross-cutting design features it exhibits and the research question(s) it primarily addresses. Feature columns: Comp./cost -compression- or cost-aware context selection; Multi-step -retrieval interleaved across reasoning or refinement steps; Safety -toxicity, bias, or privacy filtering; Personal. -user/history-conditioned retrieval; Structured -retrieval over graphs, tables, or other structure; Verify -critique or verification of retrieved evidence.
Fig. 3 : Schematic of Compression and Efficiency-Driven RAG systems.
Fig. 4 : Comparison between standard and defensive RAG pipelines, illustrating safeguards for harmful content detection, prompt sanitization, and bias-aware document retrieval.
Fig. 5 : Comparison between Normal RAG and Interactive, User-Centric RAG architectures.
Fig. 6 : Illustration of Normal RAG versus Complex Reasoning RAG. While the Naive RAG pipeline retrieves documents and generates a direct response to the initial query, the Complex Reasoning RAG decomposes the query into sub-questions, performs stepwise retrieval, and incrementally builds a structured, multi-step response tailored to complex information needs.
Method
Dataset / Setting
Accuracy / F1
Hallucination / Consistency
Retrieval Effectiveness
FoRAG
WebGPT (en), WebCPM (zh)
Factuality: 0.82–0.99
Coherence: 0.91–0.98
Training Time (Holistic): 33.1h
WeKnow-RAG
4 domains, classification, chunk-size eval
Accuracy: 0.10–0.41
Hallucination: 0.025–0.35
Confidence-aware retrieval analysis
GenGround
HotpotQA, MuSiQue, StrategyQA
F1: 27.3–52.3
Semantic Acc: 24.7–55.7
-
FRAMES
Internal multi-hop QA benchmark
Acc: 0.408–0.729
-
Prompting strategy comparison
Adobe RAG
Adobe product corpus
nDCG: 0.692–0.822
-
Dataset coverage breakdown
Self-RAG (TA-ARE)
RetrievalQA (various LLMs)
Match Acc: 6.0–46.4
-
Retrieval Acc: up to 100%
TABLE III : Performance metrics of QA-centric Retrieval-Augmented Generation models across multiple evaluation criteria. “-” indicates data not reported.
Model
Datasets / Evaluation Context
Key Metrics / Findings
PoisonedRAG
NQ, HotpotQA, MS-MARCO
ASR: 0.97–0.99; F1-Score: 0.96–1.00
RAG (CoCondenser + MiniLM)
Government, Education, Society, Health
ASR: 0.17–0.50; ASV: –0.17 to 0.67
RC-RAG
Internal eval on ChatGPT and Mistral
Risk: 14.94–19.00; Carefulness: 52.87–65.37
Towards Fair RAG
Exposure Disparity Benchmarks
EE-D: 0.14; EE-R: 0.28
LLaMA2-7B-Chat
Health, Enron Attacks
ROUGE Prompts: 73–111; Repeat Contexts: 55–135
BadRAG
GPT-4, Claude-3 Evaluation
Retrieval Success: 98.9%; Rejection Rate: 74.6%
TABLE IV : Overview of Security, Fairness, and Privacy-focused RAG Systems
Model
Category
Attribute / Metric
Score / Description
GitHub Copilot
IDE Support
Supported IDEs
IntelliJ, VSCode, PyCharm, etc.
Reference / Explanation
Provides References
No
Reference / Explanation
Explains Suggestions
No
Suggestion Variety
Options Returned
Up to 10
Training Source
Data
Public Repositories
Languages Supported
Best With
C, C++, Java, Python, etc.
TABLE V: Comparison of Retrieval-Based Code Generation Systems and Metrics
Model
Metric
Details
LLaMA2
Precision
89%
Recall
84.5%
Accuracy
85%
[1pt/1pt] RAG (Cost across prompting strategies)
Zero-shot Prompt Type I (GPT-3.5 Turbo)
$0.100
Zero-shot Prompt Type II (GPT-3.5 Turbo)
$0.014
Tree of Thoughts Prompt (GPT-3.5 Turbo)
$0.013
TABLE VI : Evaluation of RAG-based and LLM-enhanced systems in educational and corporate contexts
Retrieval-Augmented Generation (RAG) is widely regarded as a novel paradigm born from the limitations of large language models (LLMs)--a mechanism to ground their outputs in external knowledge. This view, however, is incomplete when considered within a broader historical context. In this paper, we argue that the core ideas underlying RAG are not new: foundational concepts such as integrating retrieval and language generation, knowledge augmentation, answer verification, and iterative query (or prompt) refinement had already been studied and instantiated in information retrieval (IR) and question answering (QA) research dating back to the early 2000s, well before the emergence of LLMs. We make this case by systematically tracing the intellectual lineage of modern RAG and Agentic RAG back to their classical IR and QA antecedents, and examining why this continuity has gone under-recognized -- a consequence of community fragmentation, shifting terminology, and the recency bias endemic to fast-moving fields. Rather than treating LLMs as the origin point of retrieval-augmented intelligence, we propose viewing them as a new interface layer atop a decades-old QA architecture. This reframing is not merely historical: by situating RAG within the longer trajectory of IR research, we surface underutilized prior work -- on user modeling, answer validation, and query refinement -- that can directly inform next-generation RAG design, reducing unintentional rediscovery and fostering genuine cross-community integration.
Xiaoyan Zhao, Yujie Cai, Yang Zhang +2
National University of Singapore · Georgetown University
Retrieval-Augmented Generation (RAG) enhances Large Language Models (LLMs) by grounding their responses in external knowledge, but conventional pipelines rely on static, single-step retrieval that limits performance on complex queries. This paper presents an Agent-Orchestrated Adaptive RAG framework that introduces dynamic query decomposition, iterative retrieval, and a bounded self-reflective evaluation loop. We evaluate the system across two complementary datasets: a domain-specific DevOps knowledge base and the multi-hop reasoning benchmark MuSiQue. Using metrics that include overall score, citation accuracy, mean reciprocal rank, and topic coverage, we find that query decomposition yields consistent gains in the structured domain (overall score +0.04, MRR +0.17 on DevOps) but degrades ranking precision on the multi-hop benchmark, while the reflection mechanism improves citation accuracy at a substantial latency cost. These contrasting results show that agentic enhancements are not universally beneficial and must be applied selectively according to query and domain characteristics. Our findings argue for adaptive, cost-aware orchestration rather than uniformly aggressive reasoning pipelines.
Anuj Maharjan, Devinder Kaur, Richard Molyet
Dept. of Electrical Engineering and Computer Science University of Toledo Toledo, OH, USA
Retrieval-Augmented Generation (RAG) has become a standard approach for enhancing large language models (LLMs) with external knowledge, mitigating hallucinations, and improving factuality. However, existing systems rely on generating natural language queries at each hop and maintaining a strict architectural separation between retriever and generator, preventing them from leveraging the full representational capacity of the LLM. We propose \textbf{LAnR} (Latent Abstraction for RAG), a unified framework in which a single LLM jointly performs encoding, retrieval, and generation entirely within its own latent space. Rather than generating textual queries, LAnR produces dense retrieval vectors from the hidden states of a designated \texttt{[PRED]} token and uses them to match against encoded document representations from the same model. Furthermore, LAnR adaptively decides when sufficient evidence has been retrieved using a lightweight MLP control head over those same hidden states, eliminating both the separate retriever and explicit token-level stopping reasoning. This design is motivated by our empirical observation that answer token entropy reliably signals retrieval sufficiency. Extensive experiments on six QA benchmarks spanning single-hop and multi-hop settings demonstrate that LAnR outperforms existing RAG methods, while achieving improved inference efficiency through reduced number of retrieval calls and tighter model integration.