Mitigating Hallucination in Large Language Models: A Capability-Oriented Survey on RAG, Reasoning, and Agentic Systems
Authors: Yihan Li, Xiyuan Fu, Ghanshyam Verma, Paul Buitelaar, Mingming Liu
Organizations: Electronic Information School, Wuhan University, Wuhan, China · School of Computing, Dublin City University, Dublin, Ireland · School of Public Health, Wuhan University, Wuhan, China · Insight Centre for Data Analytics, University of Galway, Ireland · Insight Centre for Data Analytics, Dublin City University, Dublin, Ireland
Hallucination remains one of the key obstacles to the reliable deployment of large language models (LLMs). Although various mitigation approaches have been proposed, existing studies often analyze different technical paradigms independently, lacking a unified perspective to understand the underlying mechanisms of different approaches and their correspondence with different types of hallucinations. This survey adopts a capability enhancement perspective to systematically examine hallucination mitigation approaches, focusing on Retrieval-Augmented Generation (RAG), reasoning enhancement, and their integration within agentic systems. Based on their primary mitigation mechanisms, we categorize hallucinations into knowledge-based hallucinations and logic-based hallucinations, analyze how RAG and reasoning enhancement methods respectively improve knowledge acquisition and reasoning reliability, and further discuss the integration mechanisms of retrieval and reasoning capabilities in Agentic Systems for mitigating composite hallucinations. By considering the applicability, mitigation mechanisms, and limitations of different approaches, this survey establishes a unified analytical framework connecting hallucination types, key capability dimensions, and technical paradigms.
Figures & tables
Figure 2. Capability-Oriented Organization of Hallucination Mitigation Techniques in LLMs A tree diagram illustrating the main content and structure of this paper, including RAG, Reasoning, Agentic System, along with their related key techniques and representative models.
Figure 3. Overview of the RAG pipeline A detailed illustration of the RAG pipeline. The diagram outlines pre-retrieval, retrieval, and post-retrieval stages, highlighting key techniques such as query rewriting, multi-turn dialogue, retrieval feedback, reranking, and knowledge integration.
Category
Key Mechanism
Representative Models
Typical Use Cases
Sparse Retrievers
Inverted index; term matching; no deep vectors
BM25 ( Lewis et al., 2020 ) , TF-IDF ( Salton and Buckley, 1988 ) , SPLADE ( Formal et al., 2021 )
Multi-hop question answering Recommender systems Intelligent reasoning
Search engine document retrieval Sources like news, academic papers, reports
Table 2. Comparison between Knowledge Graph and Unstructured Documents
Figure 6. Implementation process of three representative reasoning enhancement methods: Chain-of-Thought, Tool-augmented Reasoning and Symbolic Reasoning Illustration of different reasoning processes in large language models, including chain-of-thought reasoning, symbolic reasoning, and tool-augmented reasoning. The diagram demonstrates how step-by-step logical deduction, mathematical formulation, and external tool usage collaboratively enhance reasoning accuracy and reliability.
Figure 7. Agentic Framework Integrating RAG and Reasoning Enhancement for Comprehensive Hallucination Mitigation An overview of the general framework for hallucination mitigation in large language models, illustrating the interaction between retrieval-augmented generation (RAG), reasoning enhancement, and agentic systems. It highlights how precise and broad retrieval supply knowledge, reasoning and tool-augmented methods enhance inference, and agentic mechanisms such as reflection, planning, and memory jointly contribute to comprehensive hallucination mitigation across knowledge-based, logic-based, and composite types.
Benchmark
Hallucination Type
Data Size
Task
Evaluated Capability
Metrics
TruthfulQA ( Lin et al., 2022 )
Knowledege-based Hallucination
817
General Question Answering
Intrinsic Knowledge
Accuracy, Human Evaluation
MedHallu ( Pandit et al., 2025 )
Knowledege-based Hallucination
10,000
Medical Question Answering
RAG
F1 Score
RAGTruth ( Niu et al., 2024 )
Knowledege-based Hallucination
18000
QA, Data-to-Text Summarization
RAG
Human Evaluation
BIG-bench ( Srivastava et al., 2023 )
Logic-based Hallucination
200
Logic Reasoning
Reasoning Results
Accuracy, F1 Score
PrOntoQA ( Saparov and He, 2023 )
Logic-based Hallucination
40,000
Logic Reasoning
CoT
Accuracy
ToolBench ( Qin et al., 2024 )
Logic-based Hallucination
16,464
API invocation
Tool-Augmented Reasoning
Accuracy
Table 3. An Overview of Representative Hallucination Benchmarks
Retrieval-Augmented Generation (RAG) is widely used to augment the input to Large Language Models (LLMs) with external information, such as recent or domain-specific knowledge. Nonetheless, current models still produce closed-domain hallucinations and generate content that is unsupported by the retrieved context. Current detection approaches typically treat hallucination as a post-hoc problem, relying on black-box consistency checks or probes over frozen internal representations. In this work, we demonstrate that hallucination detection based on internal state representation can also serve as a direct training signal. We introduce RAGognize, a dataset of naturally occurring closed-domain hallucinations with token-level annotations, and RAGognizer, a hallucination-aware fine-tuning approach that integrates a lightweight detection head into an LLM, allowing for the joint optimization of language modeling and hallucination detection. This joint objective forces the model to improve the separability of its internal states regarding hallucinations while simultaneously learning to generate well-formed and meaningful responses. Across multiple benchmarks, RAGognizer achieves state-of-the-art token-level hallucination detection while substantially reducing hallucination rates during generation, without degrading language quality or relevance.
Fabian Ridder, Laurin Lessel, Malte Schilling
Computer Science Department, University of Münster, Münster, Germany
We introduce CAROL (Chain-based Adaptive Reconfiguration Over Lattices), a probabilistic framework for test-time hallucination reduction in large language models. Rather than relying on token-level uncertainty, CAROL defines a semantic uncertainty measure based on the consistency between generated responses and a trusted context, inducing a string-submodular objective over a lattice of textual sequences. This formulation enables hallucination mitigation to be cast as a Markov chain accept-reject process with provable convergence and near-optimality guarantees, allowing the model to iteratively refine outputs toward semantic consistency. By operating at the level of meaning, CAROL unifies hallucination detection and mitigation within a single framework. Empirical results on question answering and multi-agent reasoning benchmarks show that CAROL significantly reduces hallucinations and improves reliability and interpretability compared to likelihood-based and retrieval-augmented baselines, while maintaining competitive computational efficiency.
Joan Vendrell Gallart, Solmaz Kia, Russell Bent +1
Department of Mechanical and Aerospace University of California Irvine Irvine, CA 92617-4322, USA · T-5 Los Alamos National Laboratory Los Alamos, NM 88220, USA · CAI-4 Los Alamos National Laboratory Los Alamos, NM 88220, USA
Multi-step agentic retrieval-augmented generation (RAG) pipelines have demonstrated significant capability for complex reasoning tasks, yet remain vulnerable to a class of failure that existing hallucination detection mechanisms systematically miss: cascading hallucination, where errors introduced at early pipeline stages propagate and amplify across successive reasoning steps, producing confident but factually incorrect final outputs. To address this vulnerability, we formalize cascading hallucination as a distinct failure mode in agentic RAG systems, present a four-type taxonomy of cascade patterns, and introduce CHARM (Cascading Hallucination Aware Resolution and Mitigation), an architectural framework for detecting and interrupting error propagation in multi-step reasoning pipelines. CHARM comprises four components - stage-level fact verification, cross-stage consistency tracking, confidence propagation monitoring, and cascade resolution triggering - that operate alongside standard agentic RAG pipelines without requiring architectural replacement. We evaluate CHARM on HotpotQA, MuSiQue, 2WikiMultiHopQA, and a custom adversarial dataset across LangChain agentic pipeline configurations, achieving an 89.4% cascade detection rate with a 5.3% false positive rate and 215 ms +/- 18 ms average latency overhead per stage, achieving an error propagation reduction of 82.1%, compared to 18.5% for output-level detectors. Component ablations confirm that each detection module contributes meaningfully to overall cascade coverage. CHARM integrates with human-in-the-loop oversight frameworks to provide a complete reliability and governance stack for production agentic AI deployment.