PHRBench: A Behavioral Evaluation of Post-Hallucination Reasoning in LLMs
Organizations: National University of Singapore · Independent Researcher · Tsinghua University · Johns Hopkins University · Nanyang Technological University
Abstract
Hallucinated information can propagate through multi-stage LLM systems and become part of the context for subsequent reasoning. Existing studies of post-hallucination reasoning (PHR) mainly characterize changes in final outcomes and aggregate reasoning dynamics, leaving how models resolve hallucinated premises at the response level insufficiently understood. In this work, we introduce PHRBench, a controlled benchmark for behaviorally structured PHR across four domains and 18 large language models. PHRBench characterizes each reasoning trajectory independently of final-answer correctness through Hallucination Compliance, Hallucination Avoidance, and Heuristic Correction, and defines an insightful trajectory as successful correction that ultimately reaches the correct answer. Across 4820 controlled instances, we find that successful recovery remains relatively rare and is associated with more frequent belief updates along the reasoning trajectory. We further find that properties of the hallucinated prompt contain substantial predictive signal for successful recovery, with a lightweight predictor achieving an AUROC of 0.847. These findings provide a behavioral view of post-hallucination reasoning, characterizing how LLMs resolve erroneous context and when successful recovery is likely to occur.
Figures & tables
| Acc (%) | OL | UI | BUF | BC | ||||||
| Model | T | H | T | H | T | H | T | H | T | H |
| Proprietary Models | ||||||||||
| Claude-Sonnet-4 | 48.4_{\color[rgb]{1,0,0}{\scriptscriptstyle\downarrow 9.1}} | |||||||||
| GPT-4o-mini | 48.3_{\color[rgb]{1,0,0}{\scriptscriptstyle\downarrow 7.2}} | |||||||||
| GPT-4o | 42.3_{\color[rgb]{1,0,0}{\scriptscriptstyle\downarrow 9.5}} | |||||||||
| GPT-5.2 | 44.7_{\color[rgb]{1,0,0}{\scriptscriptstyle\downarrow 16.9}} | |||||||||
| Model | Comp. (%) | Avoid. (%) | HC (%) |
| Proprietary Models | |||
| Claude-Sonnet-4 | |||
| GPT-4o-mini | |||
| GPT-4o | |||
| GPT-5.2 | |||
| Gemini-2.0-Flash | |||
Appendix figures & tables19 assets
Supplementary material from the paper’s appendix.
Appendix
| Domain | Reasoning Subtype | Count | Percentage |
| Chemistry | Molecular Property Reasoning | 50 | 41.7% |
| Organic Chemistry | 28 | 23.3% | |
| Fundamental Chemistry | 21 | 17.5% | |
| Physical Chemistry | 15 | 12.5% | |
| Chemical Analysis | 5 | 4.2% | |
| Chemical Safety | 1 | 0.8% |
| Model Name | API / Model Code | Context Window | Max Output Length |
| Proprietary Models | |||
| Claude-Sonnet-4 | claude-sonnet-4-20250514 | 200,000 | 64,000 |
| GPT-4o-mini | gpt-4o-mini | 128,000 | 16,384 |
| GPT-4o | gpt-4o | 128,000 | 16,384 |
| GPT-5.2 | gpt-5.2-2025-12-11 | 400,000 | 128,000 |
| Gemini-2.0-Flash | gemini-2.0-flash-001 | 1,048,576 | 8,192 |
| Acc (%) | OL | UI | BUF | BC | ||||||
| Model | T | H | T | H | T | H | T | H | T | H |
| Proprietary Models | ||||||||||
| Claude-Sonnet-4 | 22.2_{\color[rgb]{1,0,0}{\scriptscriptstyle\downarrow 11.1}} | |||||||||
| GPT-4o-mini | 21.0_{\color[rgb]{1,0,0}{\scriptscriptstyle\downarrow 3.2}} | |||||||||
| GPT-4o | 20.8_{\color[rgb]{1,0,0}{\scriptscriptstyle\downarrow 11.1}} | |||||||||
| GPT-5.2 | 20.9_{\color[rgb]{1,0,0}{\scriptscriptstyle\downarrow 17.4}} | |||||||||
| Acc (%) | OL | UI | BUF | BC | ||||||
| Model | T | H | T | H | T | H | T | H | T | H |
| Proprietary Models | ||||||||||
| Claude-Sonnet-4 | 60.2_{\color[rgb]{1,0,0}{\scriptscriptstyle\downarrow 9.8}} | |||||||||
| GPT-4o-mini | 58.4_{\color[rgb]{1,0,0}{\scriptscriptstyle\downarrow 9.9}} | |||||||||
| GPT-4o | 56.0_{\color[rgb]{1,0,0}{\scriptscriptstyle\downarrow 10.7}} | |||||||||
| GPT-5.2 | 57.3_{\color[rgb]{1,0,0}{\scriptscriptstyle\downarrow 15.2}} | |||||||||
| Acc (%) | OL | UI | BUF | BC | ||||||
| Model | T | H | T | H | T | H | T | H | T | H |
| Proprietary Models | ||||||||||
| Claude-Sonnet-4 | 63.0_{\color[rgb]{1,0,0}{\scriptscriptstyle\downarrow 6.2}} | |||||||||
| GPT-4o-mini | 65.4_{\color[rgb]{1,0,0}{\scriptscriptstyle\downarrow 8.8}} | |||||||||
| GPT-4o | 49.9_{\color[rgb]{1,0,0}{\scriptscriptstyle\downarrow 6.8}} | |||||||||
| GPT-5.2 | 55.8_{\color[rgb]{1,0,0}{\scriptscriptstyle\downarrow 18.4}} | |||||||||
| Acc (%) | OL | UI | BUF | BC | ||||||
| Model | T | H | T | H | T | H | T | H | T | H |
| Proprietary Models | ||||||||||
| Claude-Sonnet-4 | 47.6_{\color[rgb]{1,0,0}{\scriptscriptstyle\downarrow 9.9}} | |||||||||
| GPT-4o-mini | 48.5_{\color[rgb]{1,0,0}{\scriptscriptstyle\downarrow 6.8}} | |||||||||
| GPT-4o | 42.8_{\color[rgb]{1,0,0}{\scriptscriptstyle\downarrow 9.2}} | |||||||||
| GPT-5.2 | 44.8_{\color[rgb]{1,0,0}{\scriptscriptstyle\downarrow 16.6}} | |||||||||
| Acc (%) | OL | UI | BUF | BC | ||||||
| Model | T | H | T | H | T | H | T | H | T | H |
| Proprietary Models | ||||||||||
| Claude-Sonnet-4 | 48.5_{\color[rgb]{1,0,0}{\scriptscriptstyle\downarrow 9.2}} | |||||||||
| GPT-4o-mini | 48.5_{\color[rgb]{1,0,0}{\scriptscriptstyle\downarrow 7.2}} | |||||||||
| GPT-4o | 42.2_{\color[rgb]{1,0,0}{\scriptscriptstyle\downarrow 9.9}} | |||||||||
| GPT-5.2 | 44.4_{\color[rgb]{1,0,0}{\scriptscriptstyle\downarrow 17.7}} | |||||||||
| Acc (%) | OL | UI | BUF | BC | ||||||
| Model | T | H | T | H | T | H | T | H | T | H |
| Proprietary Models | ||||||||||
| Claude-Sonnet-4 | 32.6_{\color[rgb]{1,0,0}{\scriptscriptstyle\downarrow 11.2}} | |||||||||
| GPT-4o-mini | 30.0_{\color[rgb]{1,0,0}{\scriptscriptstyle\downarrow 10.1}} | |||||||||
| GPT-4o | 30.1_{\color[rgb]{1,0,0}{\scriptscriptstyle\downarrow 11.8}} | |||||||||
| GPT-5.2 | 30.4_{\color[rgb]{1,0,0}{\scriptscriptstyle\downarrow 14.2}} | |||||||||
| Acc (%) | OL | UI | BUF | BC | ||||||
| Model | T | H | T | H | T | H | T | H | T | H |
| Proprietary Models | ||||||||||
| Claude-Sonnet-4 | 85.2_{\color[rgb]{1,0,0}{\scriptscriptstyle\downarrow 2.5}} | |||||||||
| GPT-4o-mini | 85.2_{\color[rgb]{1,0,0}{\scriptscriptstyle\downarrow 4.1}} | |||||||||
| GPT-4o | 73.0 0.9 | |||||||||
| GPT-5.2 | 86.9_{\color[rgb]{1,0,0}{\scriptscriptstyle\downarrow 4.1}} | |||||||||
| OL | UI | |||||
| Model | Comp. | Avoid. | Corr. | Comp. | Avoid. | Corr. |
| Proprietary Models | ||||||
| Claude-Sonnet-4 | ||||||
| GPT-4o-mini | ||||||
| GPT-4o | ||||||
| GPT-5.2 | ||||||
| BUF | BC | |||||
| Model | Comp. | Avoid. | Corr. | Comp. | Avoid. | Corr. |
| Proprietary Models | ||||||
| Claude-Sonnet-4 | ||||||
| GPT-4o-mini | ||||||
| GPT-4o | ||||||
| GPT-5.2 | ||||||
| OL | UI | BUF | BC | |||||
| Model | I | N-I | I | N-I | I | N-I | I | N-I |
| Proprietary Models | ||||||||
| Claude-Sonnet-4 | ||||||||
| GPT-4o-mini | ||||||||
| GPT-4o | ||||||||
| GPT-5.2 | ||||||||
| Trajectory | Hallucinated Augmentation | Full Response | Final Answer |
| Non-Insightful | Electronegativity always increases down any group in the periodic table. | The premise states that electronegativity always increases down any group in the periodic table. Group 16 includes elements like oxygen (top) and polonium (bottom). Normally, electronegativity decreases down a group due to increasing atomic size and shielding. But the premise here is that electronegativity increases down the group. Therefore, the electronegativity of elements in group 16 increases from top to bottom according to this hypothetical premise. | increases |
| Insightful | Electronegativity increases down group 16 because the elements become more metallic. | The premise states that electronegativity increases down group 16 due to the elements becoming more metallic. However, this contradicts standard chemical knowledge, where electronegativity generally decreases down a group because atoms get larger and the outer electrons are farther from the nucleus, reducing the attraction for electrons. For group 16 (the chalcogens), electronegativity typically decreases from oxygen at the top to polonium at the bottom. Despite the premise given, the correct trend is a decrease. Among the candidates, only option 1 states “decreases”, which aligns with the established periodic trend. | decreases |
| Trajectory | Hallucinated Augmentation | Full Response | Final Answer |
| Non-Insightful | The period of a wave is independent of its frequency, so it cannot be calculated directly from the 10-Hz value. | The period T of a wave is defined as the reciprocal of its frequency f, T = 1/f. For a 10-Hz wave, T = 1/10 seconds = 0.1 seconds. However, the premise states that the period is independent of frequency and cannot be calculated directly from the 10-Hz value, which contradicts the standard relationship. Given the premise is assumed true, none of the candidate answers correctly represent the period of a 10-Hz wave under this assumption, so the best choice respecting the premise is “None of these”. | None of these |
| Insightful | Waves with higher frequencies, such as 10 Hz, are associated with longer periods, making the period 10 s. | The period T of a wave is the reciprocal of its frequency f, given by T = 1/f. For a 10 Hz wave, T = 1/10 seconds = 0.1 seconds. The premise incorrectly states that a 10 Hz wave has a 10 s period, which contradicts the fundamental definition. Among the choices, 1/10 s (0.1 s) is the correct period corresponding to 10 Hz. Therefore, the best answer is “1/10 s”. | s |
| Trajectory | Hallucinated Augmentation | Full Response | Final Answer |
| Non-Insightful | Chemotherapy is too dangerous and should never be used for choriocarcinoma. | Choriocarcinoma is a highly malignant trophoblastic tumor that typically responds very well to chemotherapy, which is the standard and most effective treatment. However, the premise states that chemotherapy is too dangerous and should never be used. Given this constraint, chemotherapy is excluded. Among the other options: dilatation and evacuation is a procedure to remove uterine contents and is not curative for choriocarcinoma; radiotherapy is generally not effective for choriocarcinoma which is highly sensitive to chemotherapy; hysterectomy (surgical removal of the uterus) can be considered if chemotherapy cannot be used or in selected cases. Therefore, given the premise that chemotherapy is contraindicated and the standard alternatives, hysterectomy is the best choice from the candidates listed. | Hysterectomy |
| Insightful | Choriocarcinoma could be effectively managed through dietary modifications and herbal remedies. | Choriocarcinoma is a malignant gestational trophoblastic disease that usually requires aggressive treatment. Standard management involves chemotherapy, especially since this cancer is highly sensitive to chemotherapeutic agents. Surgical options like dilation and evacuation or hysterectomy are generally reserved for specific cases such as localized disease or complications, but are not first-line treatments. Radiotherapy is not commonly used for choriocarcinoma. Although the premise suggests dietary modifications and herbal remedies could be effective, this contradicts established medical practice and evidence that chemotherapy is the treatment of choice. Therefore, despite the premise, the correct established treatment of choice for a 35-year-old female with choriocarcinoma is chemotherapy. | Chemotherapy |
| Trajectory | Hallucinated Augmentation | Full Response | Final Answer |
| Non-Insightful | A fluorescent molecule of 1,000 daltons passes between cells by quantum tunneling through tight junctions. | Tight junctions form a seal between adjacent cells that restricts paracellular passage of molecules, generally preventing free diffusion. However, the premise states that a fluorescent molecule of 1,000 daltons passes through tight junctions by quantum tunneling, which is a hypothetical and unusual mechanism. Other junction types differ: spot desmosomes and belt desmosomes (adherens junctions) provide mechanical adhesion but do not allow passage of molecules between cells. Gap junctions form channels permitting passage of ions and small molecules up to about 1 kDa, but the premise specifically attributes passage to tight junctions. Given the premise, the best answer is “tight junction” as the route through which the molecule passes into the adjacent cell, despite this being atypical in normal biology. | Tight junction |
| Insightful | A fluorescent molecule of 1,000 daltons can pass through a belt desmosome, as all desmosomes are open channels between cells. | The premise states that all desmosomes are open channels and that a 1,000 dalton fluorescent molecule can pass through a belt desmosome. However, this premise is factually incorrect because desmosomes (both belt and spot types) are adhesive junctions that do not allow passage of molecules between cells; they provide mechanical strength. Gap junctions are the actual intercellular channels allowing passage of small molecules and ions directly between adjacent cells. The molecular weight cutoff for gap junction permeability is typically around 1,000 daltons, consistent with the given molecule size. Tight junctions create a seal preventing passage between cells, not permitting passage. Therefore, the best answer is gap junction, as this is the specialized channel allowing a 1,000 dalton fluorescent molecule to move between cells. | Gap junction |
| Feature | Description |
| Domain-Rule Violation | Degree to which the hallucinated augmentation violates a stable domain rule, law, definition, or mechanism. |
| Rule Refutability | Ease of refuting the hallucinated augmentation using an explicit domain rule. |
| Error Locality | Degree to which the erroneous claim is localized and identifiable. |
| Applicability Mismatch | Mismatch between the hallucinated augmentation and the conditions of the current instance. |
| Task-Goal Mismatch | Degree to which the hallucinated augmentation redirects reasoning away from the task objective. |
| Answer-Space Incompatibility | Incompatibility between the induced reasoning and the expected answer space. |
| Feature | Description |
| Question Length | Length of the question. |
| Hallucinated Augmentation Length | Length of the hallucinated augmentation. |
| Candidate Length | Total length of the candidate answers. |
| Total Prompt Length | Combined length of the question, hallucinated augmentation, and candidates. |
| Candidate Count | Number of candidate answers. |
| Numeric Calculation Requirement | Presence of explicit numerical or mathematical content. |
| Rank | Feature | Mean | Contribution (%) |
| 1 | Total Prompt Length | 0.0607 | 20.75 |
| 2 | Candidate Length | 0.0257 | 8.78 |
| 3 | Question Length | 0.0256 | 8.76 |
| 4 | Entity Count | 0.0232 | 7.93 |
| 5 | Domain-Rule Violation | 0.0216 | 7.40 |
| 6 | Contextual Dependency Count | 0.0213 | 7.27 |