Mitigating Hallucination in Large Language Models: A Capability-Oriented Survey on RAG, Reasoning, and Agentic Systems
Authors: Yihan Li, Xiyuan Fu, Ghanshyam Verma, Paul Buitelaar, Mingming Liu
Organizations: Electronic Information School, Wuhan University, Wuhan, China · School of Computing, Dublin City University, Dublin, Ireland · School of Public Health, Wuhan University, Wuhan, China · Insight Centre for Data Analytics, University of Galway, Ireland · Insight Centre for Data Analytics, Dublin City University, Dublin, Ireland
Hallucination remains one of the key obstacles to the reliable deployment of large language models (LLMs). Although various mitigation approaches have been proposed, existing studies often analyze different technical paradigms independently, lacking a unified perspective to understand the underlying mechanisms of different approaches and their correspondence with different types of hallucinations. This survey adopts a capability enhancement perspective to systematically examine hallucination mitigation approaches, focusing on Retrieval-Augmented Generation (RAG), reasoning enhancement, and their integration within agentic systems. Based on their primary mitigation mechanisms, we categorize hallucinations into knowledge-based hallucinations and logic-based hallucinations, analyze how RAG and reasoning enhancement methods respectively improve knowledge acquisition and reasoning reliability, and further discuss the integration mechanisms of retrieval and reasoning capabilities in Agentic Systems for mitigating composite hallucinations. By considering the applicability, mitigation mechanisms, and limitations of different approaches, this survey establishes a unified analytical framework connecting hallucination types, key capability dimensions, and technical paradigms.
Figures & tables
Figure 2. Capability-Oriented Organization of Hallucination Mitigation Techniques in LLMs A tree diagram illustrating the main content and structure of this paper, including RAG, Reasoning, Agentic System, along with their related key techniques and representative models.
Figure 3. Overview of the RAG pipeline A detailed illustration of the RAG pipeline. The diagram outlines pre-retrieval, retrieval, and post-retrieval stages, highlighting key techniques such as query rewriting, multi-turn dialogue, retrieval feedback, reranking, and knowledge integration.
Category
Key Mechanism
Representative Models
Typical Use Cases
Sparse Retrievers
Inverted index; term matching; no deep vectors
BM25 ( Lewis et al., 2020 ) , TF-IDF ( Salton and Buckley, 1988 ) , SPLADE ( Formal et al., 2021 )
Multi-hop question answering Recommender systems Intelligent reasoning
Search engine document retrieval Sources like news, academic papers, reports
Table 2. Comparison between Knowledge Graph and Unstructured Documents
Figure 6. Implementation process of three representative reasoning enhancement methods: Chain-of-Thought, Tool-augmented Reasoning and Symbolic Reasoning Illustration of different reasoning processes in large language models, including chain-of-thought reasoning, symbolic reasoning, and tool-augmented reasoning. The diagram demonstrates how step-by-step logical deduction, mathematical formulation, and external tool usage collaboratively enhance reasoning accuracy and reliability.
Figure 7. Agentic Framework Integrating RAG and Reasoning Enhancement for Comprehensive Hallucination Mitigation An overview of the general framework for hallucination mitigation in large language models, illustrating the interaction between retrieval-augmented generation (RAG), reasoning enhancement, and agentic systems. It highlights how precise and broad retrieval supply knowledge, reasoning and tool-augmented methods enhance inference, and agentic mechanisms such as reflection, planning, and memory jointly contribute to comprehensive hallucination mitigation across knowledge-based, logic-based, and composite types.
Benchmark
Hallucination Type
Data Size
Task
Evaluated Capability
Metrics
TruthfulQA ( Lin et al., 2022 )
Knowledege-based Hallucination
817
General Question Answering
Intrinsic Knowledge
Accuracy, Human Evaluation
MedHallu ( Pandit et al., 2025 )
Knowledege-based Hallucination
10,000
Medical Question Answering
RAG
F1 Score
RAGTruth ( Niu et al., 2024 )
Knowledege-based Hallucination
18000
QA, Data-to-Text Summarization
RAG
Human Evaluation
BIG-bench ( Srivastava et al., 2023 )
Logic-based Hallucination
200
Logic Reasoning
Reasoning Results
Accuracy, F1 Score
PrOntoQA ( Saparov and He, 2023 )
Logic-based Hallucination
40,000
Logic Reasoning
CoT
Accuracy
ToolBench ( Qin et al., 2024 )
Logic-based Hallucination
16,464
API invocation
Tool-Augmented Reasoning
Accuracy
Table 3. An Overview of Representative Hallucination Benchmarks
Department of Mechanical and Aerospace University of California Irvine Irvine, CA 92617-4322, USA · T-5 Los Alamos National Laboratory Los Alamos, NM 88220, USA · CAI-4 Los Alamos National Laboratory Los Alamos, NM 88220, USA