Cite What You Explore: Budget-Aware LLM Reasoning over Medical KGs with Verifiable Evidence
Authors: Chen Chen, Dongjie Wang, Mei Liu, Zijun Yao
Organizations: Electrical Engineering and Computer Science, University of Kansas, USA · Health Outcomes and Biomedical Informatics, University of Florida, USA
Post-discharge risk prediction from electronic health records (EHRs) is difficult because many dependencies that link discharge-time observations to downstream complications, such as comorbidity cascades and drug-disease interactions, are absent from the record. External medical knowledge graphs (KGs) can supply these missing dependencies, but tracing them demands three properties: KG exploration must remain cost-bounded, retrieved evidence must be differentiated by source quality, and the resulting rationale must be citable for retrospective review. Large language models (LLMs) can plan and verify over structured evidence, making them natural candidates for KG reasoning, but existing LLM-based methods do not satisfy these three properties jointly. In this paper, we propose BAR, a Budget-Aware LLM Reasoning framework over medical KGs with three contributions. First, BAR refines the raw KG into disease-specific evidence graphs whose edges carry support scores and provenance records, turning the KG into a quality-annotated reasoning space rather than a static feature source. Second, an LLM then reasons over this graph through a plan-navigate-verify loop that decomposes the question into steps, retrieves evidence under a patient-specific budget, and revises when verification fails. Third, a reasoning policy is trained with a reward that compares predictions with and without acquired evidence, combined with acquisition cost and citation-integrity terms. Across 8 diseases and 3 prediction horizons on MIMIC-III and MIMIC-IV, BAR improves AUPRC by 3.4 points over the strongest baseline, raises citation precision from 59.8% to 77.9%, and consumes only 62-65% of the budget cap.
Figures & tables
Figure 1: Overview of BAR. Stage 1: The raw KG is refined into a disease-specific evidence graph with per-edge support, provenance, and deduplication metadata. Stage 2: The LLM generates a budget-aware reasoning plan, navigates the evidence graph using a patient-conditioned quality score, and verifies each step for consistency and the cost-quality trade-off, revising as needed. Stage 3: The reasoning policy is trained via paired gain, acquisition cost, and citation-integrity reward.
Mimic-III
Mimic-IV
Category
Method
AUROC
AUPRC
AUROC
AUPRC
EHR-only
RETAIN
71.38 ±0.87
9.26 ±0.41
70.52 ±1.03
8.81 ±0.47
StageNet
72.95 ±0.74
10.53 ±0.52
72.41 ±0.81
9.67 ±0.43
Med-BERT
74.63 ±0.58
11.37 ±0.45
74.17 ±0.69
11.02 ±0.51
Static KG
GRAM
75.21 ±0.76
12.48 ±0.53
75.38 ±0.68
11.85 ±0.57
GraphCare
77.42 ±0.63
14.35 ±0.48
76.89 ±0.71
13.71 ±0.52
Table 1: Main results (%). Prediction quality is averaged across 8 diseases and 3 horizons. Reasoning quality is reported for methods with traces. † marks significance at p<0.05 vs. the strongest baseline.
Figure 2: Ablation results from the full model, averaged across both datasets. Colors indicate framework stage.
Figure 3: AUPRC and citation precision vs. budget B (left), and per-category budget utilization (right), averaged across both datasets.
Figure 4: Reasoning trace for a patient.
Appendix figures & tables11 assets
Supplementary material from the paper’s appendix.
Appendix
Cohort Size
MIMIC-III Pos. Rate (%)
MIMIC-IV Pos. Rate (%)
Target Disease
III
IV
30d
90d
365d
30d
90d
365d
Chronic kidney disease
28,143
41,672
1.83
4.21
11.47
1.71
3.98
10.83
Heart failure
25,817
38,246
2.14
5.63
13.28
2.03
5.27
12.61
COPD exacerbation
32,461
48,137
1.47
3.85
9.63
1.38
3.62
9.14
Hyperkalemia
35,284
52,318
3.26
6.14
10.82
3.08
5.81
10.24
Hypoglycemia
36,127
54,893
1.92
4.53
8.71
1.79
4.28
8.32
Appendix
Table 2: Dataset statistics. Each disease defines a separate cohort after excluding patients with pre-existing diagnoses. Splits are temporal (70/10/20).
Target Disease
Nodes
Edges
Supp. Rem. (%)
Equiv. Red. (%)
Guide Cov. (%)
Chronic kidney disease
6,841
18,273
42.3
13.7
87.4
Heart failure
7,214
19,847
41.8
14.2
89.1
COPD exacerbation
4,923
13,516
46.1
11.3
84.6
Hyperkalemia
3,847
10,284
48.7
10.8
82.3
Hypoglycemia
3,612
9,738
49.2
9.6
80.7
Ischemic stroke
5,731
16,142
43.6
12.9
86.8
Appendix
Table 3: Per-disease evidence graph after refinement from PrimeKG [ 4 ] ( ∼ 129K nodes, 4.0M edges). Comorbidity cascade diseases produce larger graphs due to more intermediate concepts.
Evidence Graph
Reasoning Loop
Hop limit K
3
Budget Bmin / Bmax
2 / 12
Degree cap dmax
50
Query / expand / select cost
1.0 / 0.5 / 0.5
Hub edge cap M
10
Pre-filter cap Kpre
5
Support threshold τs
0.3
Citation cap KC
5
Min edges Kmin
2
Stop ϵstop / patience Pstop
0.01 / 2
Quality threshold τfb
0.4
Evidence cap KE
8
Appendix
Table 4: Hyperparameters and configuration.
Time Breakdown
Wall-Clock (s)
Method
LLM Calls
Graph (ms)
Scoring (ms)
Verify (ms)
Steps
Mean
P95
ToG-adapt
13.4 ±1.2
84 ±12
31 ±5
—
6.7 ±0.4
8.7 ±1.1
12.4
VoG-adapt
11.2 ±0.9
72 ±10
28 ±4
48 ±7
5.6 ±0.3
7.3 ±0.8
10.1
BAR (GPT-4o-mini)
7.8 ±0.6
53 ±8
42 ±6
38 ±5
3.9 ±0.3
5.1 ±0.7
7.8
BAR (Qwen3-8B, local)
7.8 ±0.6
53 ±8
42 ±6
38 ±5
3.9 ±0.3
2.3 ±0.4
3.6
Appendix
Table 5: Inference latency per patient on MIMIC-III test set (single L40S GPU, GPT-4o-mini API). BAR is fastest because fewer reasoning steps mean fewer LLM calls.
Epoch 1
Epoch 10
Epoch 20
Mean episode reward
0.24
0.41
0.44
Reward std
0.34
0.22
0.19
Val AUPRC (%)
17.84
19.12
19.57
RL gain : Δ AUPRC =+1.73±0.47 (stage-2 → stage-3, across 5 seeds)
Appendix
Table 6: Stage-3 training dynamics. Reward statistics are averaged across training batches at the indicated epoch. The AUPRC gain from RL is consistent across 5 random seeds.
Method
CKD
HF
COPD
HyperK
HypoG
Stroke
Sepsis
DVT
Avg
Mimic-III
RETAIN
10.83 ±0.81
11.42 ±0.78
8.71 ±0.94
12.17 ±0.74
8.34 ±0.91
5.63 ±1.34
11.28 ±0.83
5.72 ±1.08
9.26 ±0.41
StageNet
12.14 ±0.79
12.87 ±0.76
9.83 ±0.92
13.52 ±0.72
9.21 ±0.89
6.42 ±1.31
12.63 ±0.81
7.61 ±1.06
10.53 ±0.52
Med-BERT
13.28 ±0.83
13.94 ±0.81
10.62 ±0.96
14.38 ±0.76
10.17 ±0.93
7.14 ±1.36
13.41 ±0.85
8.02 ±1.10
11.37 ±0.45
GRAM
14.52 ±0.86
15.31 ±0.83
11.47 ±0.99
15.63 ±0.79
10.84 ±0.96
7.83 ±1.39
14.27 ±0.88
9.97 ±1.13
12.48 ±0.53
GraphCare
16.47 ±0.91
17.28 ±0.88
13.14 ±1.04
17.42 ±0.84
12.53 ±1.01
9.41 ±1.46
16.18 ±0.93
12.37 ±1.18
14.35 ±0.48
ToG-adapt
18.36 ±1.05
19.24 ±1.01
14.47 ±1.18
19.73 ±0.96
14.28 ±1.14
10.87 ±1.62
18.12 ±1.06
14.37 ±1.32
16.18 ±0.63
Appendix
Table 7: AUPRC (%) per disease, mean ± std over 5 seeds. Best in bold among distinct methods.
Figure 5: Per-horizon AUPRC averaged across both datasets. Percentages show the relative gain of BAR over the best EHR-only baseline.
Figure 6: Robustness analysis on MIMIC-III. (a) AUPRC under random ICD code dropping at test time (no retraining). (b) Disease transfer via leave-2-out evaluation: train the full pipeline on 6 diseases, fine-tune only prediction heads on 2 held-out diseases (all other parameters frozen). Percentages show AUPRC decline on held-out vs. seen diseases.
Figure 7: Sensitivity to key hyperparameters, averaged across 2 datasets. Gray bands mark defaults.
High Relevance
Low Relevance
Marginal Rate
High Support
68.4
31.2
49.8
Low Support
22.7
14.5
18.6
Marginal Rate
45.6
22.9
—
Appendix
Table 8: Selection rate (%) into the final evidence set, partitioned by median-split quadrants of the support score and patient-conditioned relevance term. Edges must be strong on both dimensions to be selected at high rates, validating the multiplicative design.
Annotator A
Annotator B
GPT-4o
Annotator A
—
0.74
0.82
Annotator B
0.74
—
0.69
GPT-4o
0.82
0.69
—
vs. other-two consensus
0.81
0.71
0.78
Appendix
Table 9: Pairwise inter-rater agreement (Cohen’s κ ) on 300 patient-edge pairs. The bottom row reports agreement between each individual rater and the majority-vote consensus of the other two.
Large language models (LLMs) can generate biomedical hypotheses, but it remains unclear whether they truly reason from scientific evidence or simply produce convincing-sounding ideas. To study this, we combine three major biological databases: the Kyoto Encyclopedia of Genes and Genomes (KEGG), Rhea, and UniProt, into a unified biochemical knowledge graph and construct a benchmark of 550 paths connecting enzyme sources to rare disease endpoints, yielding 13,200 hypotheses from six LLMs under four conditions varying the biological information each model receives: source enzyme only, full biological path, or source and disease endpoint only. Hypotheses are scored using an expert-derived five-criterion rubric on a 1-5 scale per criterion. We find that models given both the source and disease endpoint often produce the highest-scoring hypotheses, showing that LLMs can generate compelling ideas from minimal information. However, these hypotheses are less grounded in the evidence. In contrast, models given the full biological path generate hypotheses more consistent with known mechanistic relationships. We call this evidence-disciplined reasoning. To confirm this effect, we shuffled intermediate path steps while keeping endpoints fixed. Evidence grounding dropped significantly (delta = -0.793, p < 0.001), confirming models genuinely used path structure during reasoning. Our findings show that knowledge graphs support hypothesis generation in two ways: they identify biological endpoint pairs absent from the literature, and their mechanistic paths guide how LLMs reason between them.
Dominic Okonkwo, Adetayo Okunoye, Ismailcem Budak Arpinar
Faithful reasoning is essential in medicine, where clinical decisions require transparent justification grounded in reliable evidence. Current medical LLMs either lack active access to evidence or use retrieved evidence without supervising how it should be appraised and applied during reasoning. To address this, we formalize evidence-based medicine principles as process-level criteria and introduce FaithMed, a framework that combines clinician-designed, automatically refined rubrics with reinforcement learning using step-level process reward assignment and advantage grouping. Across seven medical benchmarks, FaithMed improves over agentic-search baselines (+9% on average) and outcome-only RL (+5.8%), while raising average evidence-based medicine rubric scores over agentic-search Qwen3 baselines (+15.5%). This work demonstrates that explicit step-level supervision can improve both task success and the faithfulness of the reasoning process. Code is available at https://github.com/cxcscmu/FaithMed.
Zhiyun Zhang, Liwen Sun, Xiang Qian +1
Carnegie Mellon University · Stanford University School of Medicine · Xlue
Longitudinal electronic health records (EHRs) capture years of patient history across notes, codes, labs, and procedures, and contain evidence needed to reason about likely clinical outcomes. However, comprehensive clinician review of these records is impractical, and LLM-based processing is costly and often unreliable, missing some relevant observations while hallucinating others. We therefore propose EviGen, a three-layer framework for verifiable clinical rationale generation that addresses these challenges. The first layer is a patient-conditioned retriever that uses learnable queries to find evidence predictive of, not just textually relevant to, a clinical outcome and ranks it by prediction attribution scores. The second layer is an LLM generator that consumes this ranked evidence as a scaffold to produce a clinical rationale grounded in the retrieved spans. The third layer is a process-supervised verifier that checks the generated rationale at the reasoning-step level, flagging unreliable claims. Across three medical prediction datasets, EviGen improves prediction performance and rationale faithfulness over full-context LLM and RAG baselines, and is preferred by clinical reviewers in a usability evaluation.