Hallucination Across the Reasoning Lifecycle: Interface Visibility, Causal Evidence, and Release Control in Large Reasoning Models
Organizations: Zhejiang University, China · Binjiang Institute of Zhejiang University, China · Guangzhou University, China · Rensselaer Polytechnic Institute, USA · The University of Hong Kong, Hong Kong SAR, China · The Hong Kong University of Science and Technology (Guangzhou), China · The Hong Kong Polytechnic University, Hong Kong SAR, China · Research Center for Scientific Data Hub, Zhejiang Lab, China
Abstract
Reasoning errors can propagate into later decisions and memory. This survey synthesizes 312 papers and first-party reports on text-based reasoning hallucinations around three questions: what evidence is observable, what study designs establish, and which corrective actions the evidence supports. UIPCA records unsupported premises (U), invalid inferences (I), dependent reuse (P), visible answer-trace consistency (C), and action-policy failures (A). Across 58 reviewed sources, no comparison establishes that a specified intervention improves reasoning while reducing factual reliability under matched conditions. The synthesis connects diagnosis to verification, repair, selective release, and persistent-state control across memory, tools, and training feedback.
Figures & tables
| Survey / Focus | Interface and Evidence Design | Persistence and Action | What This Review Adds |
|---|---|---|---|
| Broad and domain hallucination surveys ( Ji et al., 2023 ; Huang et al., 2025 ; Zhang et al., 2025 ; Sahoo et al., 2024 ; Lee et al., 2025 ) | Organize definitions, causes, benchmarks, mitigation, or domain-specific errors; interface disclosure is not the main axis | Controls are usually attached to an output, model stage, or domain failure | Retain their error types and checks, then record visibility, evidence design, dependency reuse, and action |
| Ho et al. ( Ho et al., 2026 ) | Multi-level definition and model lifecycle connect text and vision; causal matching is not the main organizing idea | Causes and mitigations span model families, but released claims are not tracked through dependencies or later writes | Add interface-specific unknown labels, study-design classes, and claim reuse |
| Song et al. ( Song et al., 2026 ) | Crosses reasoning types with fundamental, application, and robustness failures | Identifies causes and mitigation directions across informal, formal, and embodied reasoning | Add error location and separate process diagnosis from release-policy failure |
| Long- and image-grounded CoT surveys ( Chen et al., 2026c ; Dong et al., 2026 ) | Compare trace length, training method, task demand, or interleaved states | Emphasize capability and method trade-offs rather than later unsupported writes | Use their method distinctions while checking support, visibility, and action |
| Liu et al. ( Liu et al., 2026 ) | Evaluates reasoning hallucination in an educational setting and derives a task-specific classification | Tests uncertainty-based abstention and model-to-model correction within that application | Supply the broader, source-level evidence synthesis and interface-aware diagnostic record |
| Mirages of Logic ( Anonymous ACL submission, 2026 ) | Types visible premise, operation, logic, and conclusion failures | Supports local process diagnosis; dependency reuse, interface opacity, and release action are outside its core type labels | Retain error type, then add origin/reuse, unknown-by-interface, answer–trace consistency, and policy action |
| Evaluation Target | Operational Question | Typical Unit / Evidence | Interface Boundary |
|---|---|---|---|
| Factuality / source grounding | Is a claim true at evaluation time, or licensed by the designated source? | Atomic claim or answer; fact label, source entailment, execution result | Available under all interfaces when the required external evidence exists |
| Process faithfulness | Does the displayed reasoning reflect what causally determines the answer? | Trace–answer pairs with controlled perturbations ( Lanham et al., 2023 ; Tutek et al., 2025 ) | Causal attribution requires intervention; textual consistency alone is insufficient |
| Uncertainty calibration | Do reported probabilities or scores match empirical correctness? | Answer set; Brier, ECE, calibration curve ( Damani et al., 2025 ; Khanmohammadi et al., 2026 ) | Behavioral confidence can be evaluated under all interfaces; internal signals require provider access |
| Corrigibility | Does valid feedback change the dependent trace and outcome? | Edited step or trajectory ( Lu et al., 2025 ) | Requires before/after evidence; local acceptance is not outcome recovery |
| Selective decision quality | Does a fixed policy choose answer, retrieve, verify, clarify, abstain, or escalate appropriately? | Item and action; selective risk, coverage, action counts | The action is observable; its justification depends on declared policy inputs and evidence |
| Setting | Premise Object | Early Check | Later Check / Action | Unresolved Boundary |
|---|---|---|---|---|
| Source-grounded text | Passage, database record, or time-indexed fact | Claim extraction and source entailment ( Chern et al., 2023 ; Hu et al., 2024 ) | Verify, retrieve, revise, abstain | Hidden evidence use and open-world completeness |
| Image or video | Object, region binding, frame, or temporal segment | Coverage and grounding ( Li et al., 2023a ; Guan et al., 2024 ; Lv et al., 2026 ; Ma, 2026 ) | Reasoning check, targeted retrieval, abstention | Whether the rationale causally used the grounded evidence |
| Function or repository code | Requirement, API, dependency, program state, repository snapshot | Type/contract/static check ( Tian et al., 2024 ; Erfanian et al., 2026 ) | Sandbox, properties, hidden tests, patch review ( Liu et al., 2023 ; Jimenez et al., 2023 ) | Untested paths, security, and unwritten intent |
| Comparison | Outcome | Outcome | Study Design | Unmet Criterion |
|---|---|---|---|---|
| Post-training recipes ( Yao et al., 2025 ) | No independent threshold | Factual QA rises or falls by recipe | Comparison of 12 separately trained systems | Data, reward, and optimizer vary together |
| System comparisons ( Chen et al., 2025a ) | System type, not matched | Higher hallucination for two reasoning systems | Reasoning versus non-reasoning systems | Checkpoint and output detail differ |
| Natural trace groups ( Gekhman et al., 2026 ) | Recall expansion reported separately | Lower accuracy with an incorrect extracted fact | Observed groups on the same questions | Facts are grouped, not experimentally manipulated |
| Qwen3 thinking mode ( Wang et al., 2026b ) | No independent | Aggregate gain with some correct-to-incorrect flips | Fixed checkpoint and items; mode varies | Prompts/decoding differ; beneficial flips coexist |
| Test-time budget ( Zhao et al., 2025 ) | Tokens, not independent capability | Non-monotonic accuracy and hallucination | Model/task fixed in many budget comparisons | Joint effect unresolved; answer willingness mediates outcomes |
| Factuality-aware RL ( Li and Ng, 2025 ) | Math reasoning accuracy | Factuality/hallucination measured separately | Training-pipeline comparison | Reasoning and factuality training vary together |
| Hypothesis | Evidence grade | Representative evidence | Decisive test still missing |
|---|---|---|---|
| H1: Unchecked claims accumulate | Limited or mixed | Faulty traces and complexity-related errors ( Lu et al., 2025 ; Heyman and Zylberberg, 2025 ) | Fix the initial trace; vary unchecked claims, verification, and stopping |
| H2: Rewards favor answering | Supported synthesis; bounded reward interventions | Reward granularity and calibrated reflection ( Palandye et al., 2026 ; Damani et al., 2025 ) | Randomize reward arms from one base checkpoint; fix data, budget, optimization, and evaluation; use multiple seeds |
| H3: Relevance outweighs evidence | Limited or mixed | Wrong guidance and sensitivity to supplied evidence ( Lu et al., 2025 ; Shi et al., 2023 ) | Vary relevance and support for one task using matched wording templates |
| H4: Errors are reused | Direct, single-setting seeded attack | FARMA: forged reasoning stored and reintroduced ( Karamchandani et al., 2026 ) | Seed supported, unsupported, and refuted premises; vary independent repair and measure actual reuse |
| H5: Uncertainty changes action | Supported synthesis; limited transfer | Uncertainty-guided control and the monitoring–control gap ( Yan et al., 2026 ; Yu et al., 2026b ) | Change only the signal; measure action and useful coverage, then test distribution shift |
| Signal / Examples | Unit / Covered Component | Interface / Access | Feasible Action | Boundary and Cost |
|---|---|---|---|---|
| External evidence: CLATTER, RefChecker ( Eliav et al., 2025 ; Hu et al., 2024 ) | Claim; U and source-backed I | All interfaces; source required | Accept, retrieve, revise, abstain | Strong support; extraction and checking cost grows with the number of claims |
| Uncertainty or learned scoring: Semantic Entropy Probes, FaithLens, UPAIR ( Kossen et al., 2024 ; Si et al., 2025 ; Yan et al., 2026 ) | Answer/trace; routes checks rather than establishing support or C | White-box probe or learned detector | Verify, resample, reject | Checkpoint and training-domain dependence |
| Verification and calibrated action: VeriFY, CURE, Latent Critic ( Altinisik et al., 2026 ; Liu and Wang, 2026 ; Vijayvargiya and Lokesh, 2026 ) | Claim/answer; support checks and action selection | Black-box model, labels, or internal critic | Answer, verify, abstain | Coverage depends on boundary estimation and detector transfer |
| White-box trajectory: TRACED, HARP, Reasoning Denoiser ( Jiang et al., 2026 ; Hu et al., 2025a ; Fang et al., 2026 ) | Step/trajectory; I and P routing | Internal states or trained features | Stop, inspect, escalate | Early but checkpoint-specific; no external factual evidence |
| Online search/repair: HaluSearch, Token-Guard ( Cheng et al., 2025b ; Zhu et al., 2026 ) | Prefix/segment; I and P | Generation-time score and access | Redirect, prune, regenerate | Detection and corrective policy effects are coupled |
| Executable and grounded checks ( Ling et al., 2023 ; Pan et al., 2023 ; Ma, 2026 ; Jin et al., 2026 ) | Constraint, region, program state; U/I | Evidence, tools, or sandbox | Bind, execute, repair, abstain | The check covers only declared evidence and tested properties |
| Supervision or verifier | Target | Useful intervention | Characteristic failure |
|---|---|---|---|
| Evidence-backed premise check | Reject or retrieve before a premise enters | Open-world evidence gaps; judge shares generator bias | |
| Rule, math, or execution check | Reject an invalid local transition | Incomplete checker; valid alternative paths scored as wrong | |
| Dependency-aware reward | Stop reuse or invalidate a branch | Missing edges; rewards favor style or late correction rather than preventing reuse | |
| Answer–trace or action verifier | Compare the answer with the trace; revise, verify, abstain, or escalate | Late detection; exploitable scores, unnecessary refusal, or an unrepaired trace |
| Family | Primary Unit | Directly Reveals | Required Companion Report |
|---|---|---|---|
| Factual QA / grounding: SimpleQA, TriviaQA, FACTS ( Wei et al., 2024a ; Joshi et al., 2017 ; Cheng et al., 2025a ) | Answer or claim–source pair | Outcome factuality or source support | Evidence/evaluator version; extraction rule; answer rate and coverage |
| Trace / perturbation: FINE-CoT, C2-Faith, RACE ( Shen et al., 2026 ; Mittal and Arike, 2026 ; Wang et al., 2025a ) | Step, dependency, trace, answer | Selected process errors; coverage depends on suite labels | Segmentation, codebook, boundary cases, interface, unobserved steps |
| Calibration / answerability: RMCB, VeriFY, AbstentionBench ( Khanmohammadi et al., 2026 ; Altinisik et al., 2026 ; Kirichenko et al., 2025 ) | Confidence, answer, action | Ranking, calibration, answer/abstain behavior | Brier/ECE, selective risk, coverage, unnecessary refusal |
| Long context / changing evidence: RULER, Lost in the Middle, StreamingQA, HalluDial ( Hsieh et al., 2024 ; Liu et al., 2024 ; Liška et al., 2022 ; Luo et al., 2024 ) | Retrieved answer or turn | Outcome under available, positioned, or changing evidence | Retrieval/binding record, evidence time, first-error annotation |
| Multimodal: POPE, HallusionBench, AMBER ( Li et al., 2023a ; Guan et al., 2024 ; Wang et al., 2023 ) | Object, region, claim, answer | Perception/grounding and selected outcome | Coverage, binding, reasoning, answer, abstention as separate fields |
| Code / software: CodeHalu, CRUXEval, EvalPlus, SWE-bench ( Tian et al., 2024 ; Gu et al., 2024 ; Liu et al., 2023 ; Jimenez et al., 2023 ) | Program state, function, repair, patch | Selected execution, repair, or issue-resolution behavior | Environment, tests/coverage, analyzer, security, repository snapshot |
| System | Control point | Check or evidence | Remaining boundary |
|---|---|---|---|
| A-MAC ( Zhang et al., 2026 ) | Candidate admission and conflict-aware replacement | Weighted utility, overlap-based confidence, novelty, recency, and type prior | Admission precision/recall does not establish semantic truth; overlap is a support proxy |
| TrustMem ( Yang et al., 2026 ) | Learned write, revise, and prune behavior | Transition verifier scores coverage, preservation, and faithfulness for reward and preference training | Frozen-LLM judgments depend on the supplied chunk and prior memory; no independent truth guarantee |
| Governed Memory ( Taheri, 2026 ) | Typed extraction, entity-scoped retrieval, policy routing | Heuristic quality gates, schema checks, provenance metadata, and bounded retrieval | Self-evaluation; sequential conflict tests leave concurrent multi-agent writes unresolved |
| MemClaw / ArgusFleet ( Margalit et al., 2026 ) | Scoped access, supersession, and controlled propagation | Service traces test authorized visibility, conflict handling, and provenance reconstruction | One-service self-evaluation; admission can pre-empt correction; visibility is not dependent reasoning use |