GLoC-EHR: Evidence-Cited Clinical Reasoning over Global Context and Local EHR Events
Authors: Chaiho Shin, Kwangsoo Kim
Organizations: Interdisciplinary Program of Medical Informatics, Seoul National University College of Medicine · Department of Transdisciplinary Medicine, Seoul National University Hospital · Center for Data Science, Healthcare AI Research Institute, Seoul National University Hospital · Department of Medicine, Seoul National University College of Medicine
Structured electronic health records (EHRs) contain a patient's clinical trajectory as a sequence of clinical codes. Answering clinical questions from such records requires both the context of the whole trajectory and the specific events that support the answer. We introduce GLoC-EHR, a multimodal language model that reads a contextual encoding of the record through a fixed-size global memory of the trajectory and a local memory of selected events. The model learns to generate hospital-course summaries from the global memory and descriptions of masked concepts from the local memory, aligning both with clinical text. It is then trained to cite evidence before answering, through rationale fine-tuning followed by group relative policy optimization (GRPO) with rewards for correct answers and record-supported evidence. On three MIMIC-IV outcome tasks, GLoC-EHR attains the highest macro AUROC among the compared models when it answers directly, whereas zero-shot LLMs reading the serialized record fall far behind. With evidence-cited reasoning, it stays close to its direct multi-task counterpart in macro AUROC, and the evidence terms of the objective reduce unsupported evidence at a similar macro AUROC. The local memory adds distinct supported findings, particularly under strict matching, without a detectable change in macro AUROC. Without retraining, GLoC-EHR transfers to EHRSHOT on par with EHR-BERT and answers two unseen laboratory questions better than zero-shot prompting of its own backbone.
Figures & tables
Figure 1: Overview. Top: An EHR encoder contextualizes the observed events. The global route compresses them into a few tokens that fill the <EMR> slot and are re-read by gated cross-attention in the middle layers; the local route supplies the top- k events through gated cross-attention in the late layers. The LLM generates cited evidence, a rationale, and an answer. Bottom: \scriptsize1⃝ encoder pretraining; \scriptsize2⃝ trajectory (global) then concept (local) alignment; \scriptsize3⃝ SFT on teacher rationales followed by GRPO, whose evidence credit is scored against the patient’s record (dotted line).
AUROC ↑
Macro (3 tasks) ↑
Method
Mortality
Long LOS
Readmission
AUROC
AUPRC
(24 h)
(24 h)
(full)
Zero-shot LLMs on serialized records
Qwen3-1.7B (direct)
0.8140
0.6062
0.5183
0.6461
0.2240
Qwen3-1.7B (CoT)
0.5488
0.5363
0.5052
0.5301
0.1631
Llama-3.3-70B (direct)
0.8268
0.6391
0.5269
0.6643
0.2065
Table 1: Internal outcome prediction on MIMIC-IV (test): mean ± SD over five seeds, or over five sampled chains of one run for GLoC-EHR (Reasoning). Best in bold, second underlined. C : code counts; T4096 : text embedding; ECLS : frozen EHR-BERT vector; / LGBM: LightGBM head.
Figure 2: Evidence of the final checkpoints on MIMIC-IV test chains (one per case) as matching becomes stricter: every match must reach an IDF-weighted overlap of at least τ ( τ=0.55 is the reward’s matcher). (a) Share of supported bullets, (b) distinct supported findings per response, and (c) relative gain of GLoC-EHR (Reasoning) over the global-route-only variant, with 95% paired case-bootstrap intervals.
Macro
Macro
Unsupported
Distinct supported ↑
Bullets-only
Variant
AUROC ↑
AUPRC ↑
rate ↓
τ=0.55
τ=0.85
AUROC ↑
GLoC-EHR (Reasoning)
0.8656±0.0008
0.4819±0.0045
0.440±0.001
4.16±0.01
2.73±0.01
0.758±0.005
one chain
0.8661
0.4810
0.440
4.15
2.73
0.755
w/o evidence terms
0.8665
0.4896
0.577
3.66
2.21
0.707
global route only
0.8641
0.4783
0.431
4.00
2.41
0.767
rationale SFT only
0.6578
0.2013
0.582
3.92
2.34
0.690
Table 2: Ablations of GLoC-EHR (Reasoning) on the MIMIC-IV test set. The first row is the mean ± SD over five sampled chains; the other rows score one chain per case. Evidence is matched with the reward’s matcher ( τ=0.55 ) and, for distinct supported findings, also at τ=0.85 (Figure 2 ). Bullets-only AUROC: answers from each model’s own bullets with the EHR memories off.
AUROC ↑
Macro ↑
Method
Mortality
Long LOS
Readmission
AUROC
AUPRC
EHR-BERT (MT)
0.9311±0.0055
0.7531±0.0056
0.6314±0.0153
0.7719±0.0041
0.4414±0.0040
code IDs only
0.9303±0.0037
0.7525±0.0063
0.5568±0.0328
0.7465±0.0130
0.4167±0.0148
GLoC-EHR (MT)
0.9287±0.0023
0.7584±0.0043
0.6396±0.0131
0.7756±0.0039
0.4331±0.0060
GLoC-EHR (Reasoning)
0.9244
0.7585
0.6288
0.7706
0.4222
Table 3: External validation on EHRSHOT without retraining, with out-of-vocabulary events removed: mean ± SD over five seeds. GLoC-EHR (Reasoning): one sampled chain per case (internal five-chain SD 0.0008), excluding the 2.6% without an answer (0.767 macro AUROC if scored as 0.5). Code IDs only: EHR-BERT pretrained without description embeddings. Best in bold, second underlined.
Appendix figures & tables15 assets
Supplementary material from the paper’s appendix.
Appendix
Isotropy
Qualified concepts
Representation
Post-processing
reff↑
PC 1 % ↓
cˉ↓
cˉsame
cˉrand↓
NN c % ↑
NN q % ↓
Input-token mean
none
295
15.3
0.29
0.97
0.61
75.3
48.4
Input-token mean
whiten +ℓ2
1,793
0.2
0.00
0.94
0.00
91.1
36.4
Final hidden state
none
70
21.3
0.89
—
—
—
—
Final hidden state
top-PC removal +ℓ2
116
23.4
0.00
0.99
0.36
87.0
39.5
Final hidden state
whiten +ℓ2 (used)
1,908
0.3
0.00
0.94
0.00
98.4
31.9
Appendix
Table 4: Geometry of Qwen3-1.7B description embeddings for the 20,561 vocabulary concepts (typed short descriptions, mean pooling, 2,048 dimensions). Isotropy: effective rank reff , variance share of the first principal component, and mean pairwise cosine cˉ , which is zero by construction after centring. Qualified concepts: the 6,952 concepts with a value or exposure qualifier; cˉsame and cˉrand are mean cosines between variants of the same concept and between random qualified concepts; NN c is the share of nearest neighbours that are a variant of the same concept, among the 68.5% of qualified concepts that have one; NN q is the share whose nearest neighbour is a different concept with the same qualifier. Without the type prefix, the used embeddings give an NN c of 97.1%, and float32 inference changes NN c by less than 0.2 points.
Stage
Supervision
Trainable components
Semantic encoder
masked concepts
EHR encoder and input projections
Global alignment
BHC generation
resampler, global cross-attention
Local alignment
concept descriptions + BHC replay
projection, selector, local cross-attention; local cross-attention only after hard selection
Direct prediction
binary outcome
global and local cross-attention
Rationale SFT, soft
teacher rationale + answer
base LLM, resampler, global and local cross-attention, local projection, selector
Rationale SFT, hard
teacher rationale + answer
base LLM, resampler, global and local cross-attention
Appendix
Table 5: Optimization stages. Global and local cross-attention include their scalar gates.
Validation target
Kℓ
Full
Local off
Global off
Both off
Concept description
256
0.2797
3.3105 ( + 1083.5%)
0.5378 ( + 92.3%)
3.6285
128
0.2844
↑
0.5536 ( + 94.7%)
↑
64
0.3010
↑
0.5913 ( + 96.5%)
↑
BHC generation
256
1.7937
1.7922 ( − 0.1%)
3.2044 ( + 78.6%)
—
64
1.7942
↑
—
—
Appendix
Table 6: Route ablation after concept alignment (final Stage-2 checkpoint, hard selection): validation loss with the local and/or global cross-attention gates forced shut, and its change relative to the full model at the same Kℓ . Passes with the local route off do not read the local memory, so their values are shared across Kℓ ( ↑ : same as the Kℓ=256 row); —: not evaluated.
Eligible events
Selection inactive ( ≤Kℓ events)
Task
Median
p90
Kℓ=256
Kℓ=128
Kℓ=64
Mortality, long LOS (24 h)
96.5
603
78.9%
66.5%
29.6%
Readmission (full)
194
2,048
58.5%
37.8%
20.3%
Appendix
Table 7: Local budget on the downstream test inputs (6,804 24-h and 7,566 readmission records; mortality and long length of stay share the 24-h inputs): eligible events per record (non-padding positions of the encoder input, truncated at 2,048, a limit that 10.2% of readmission records reach) and the share of records in which top- Kℓ selection is inactive because the record has at most Kℓ eligible events. On average, Kℓ=256/128/64 selects 120.1/88.9/56.2 events for the 24-h inputs and 169.1/103.9/58.8 for readmission.
Mortality
Long LOS
Readmission
All
Cases (positives)
453 (17)
441 (152)
1,168 (75)
2,062 (244)
Persistence, cited mask
0.874
0.885
0.851
0.863
Persistence, control mask
0.885
0.889
0.875
0.880
Persistence, cited − control
−0.011 [ −0.033 , 0.010 ]
−0.005 [ −0.026 , 0.018 ]
−0.024 [ −0.040 , −0.009 ]
−0.017 [ −0.028 , −0.007 ]
∣ΔpYes∣ , cited / control
0.007 / 0.007
0.021 / 0.021
0.012 / 0.012
0.013 / 0.013
Label flips (%), cited / control
1.1 / 1.3
1.8 / 2.9
0.4 / 0.3
0.9 / 1.1
Appendix
Table 8: Cited-slot masking on MIMIC-IV test cases whose greedy chain cites concepts held in local memory. Persistence: share of the masked concepts (cited in the full chain) that the regenerated rationale still cites. ∣ΔpYes∣ and flips: change of the answer probability and of the predicted label relative to the full chain. AUROC: cited minus control, with the answer read under the mask or with the local route switched off. Brackets: two-sided 95% paired bootstrap intervals. AUROC refers to this greedy subset and is not comparable with Table 1 .
Figure 3: Cohort construction. (a) Internal MIMIC-IV cohort (OMOP CDM format), split by patient. (b) EHRSHOT external cohort; inpatient visits that start within one day of the previous discharge are merged into one admission. Task-level sample sizes are given in Table 9 .
MIMIC-IV
Task
Input
Training
Validation
Test
EHRSHOT
In-hospital mortality
First 24 h
24,876 (3.19%)
3,549 (3.18%)
6,804 (3.15%)
9,922 (2.75%)
Long LOS ( > 7 d)
First 24 h
24,876 (37.06%)
3,549 (37.70%)
6,804 (36.67%)
9,922 (31.14%)
30-day readmission
Full admission
27,420 (7.55%)
3,890 (7.46%)
7,566 (6.89%)
12,342 (23.70%)
Appendix
Table 9: Downstream sample sizes (admissions) and outcome prevalence. Mortality and long length of stay (LOS) share the same admissions and 24-h inputs. EHRSHOT is used for evaluation only.
Method
Mortality (24 h)
Long LOS (24 h)
Readmission (full)
Macro (3 tasks)
Zero-shot LLMs on serialized records (single run)
Qwen3-1.7B (direct)
0.8140
0.6062
0.5183
0.6461
Qwen3-1.7B (CoT)
0.5488
0.5363
0.5052
0.5301
Qwen3-1.7B (CoT, our prompt)
0.5203
0.5136
0.4923
0.5087
Llama-3.3-70B (direct)
0.8268
0.6391
0.5269
0.6643
Llama-3.3-70B (CoT)
0.6780
0.5193
0.5034
0.5669
Appendix
Table 10: Full internal AUROC comparison on MIMIC-IV (test): mean ± SD over five seeds, or over five sampled chains of one run for GLoC-EHR (Reasoning); rows without SD are single runs or deterministic (LR). Best in bold, second underlined. ST: one model per task; MT: one model for all three tasks. C : code counts; T4096 : text embedding; ECLS / Emean : frozen EHR-BERT vectors; a slash names the classifier head.
Method
Mortality (24 h)
Long LOS (24 h)
Readmission (full)
Macro (3 tasks)
Zero-shot LLMs on serialized records (single run)
Qwen3-1.7B (direct)
0.1203
0.4797
0.0718
0.2240
Qwen3-1.7B (CoT)
0.0347
0.3849
0.0696
0.1631
Qwen3-1.7B (CoT, our prompt)
0.0327
0.3732
0.0679
0.1579
Llama-3.3-70B (direct)
0.0916
0.4552
0.0727
0.2065
Llama-3.3-70B (CoT)
0.0480
0.3759
0.0693
0.1644
Appendix
Table 11: Full internal AUPRC comparison on MIMIC-IV (test; prevalence 3.15% / 36.67% / 6.89%): mean ± SD over five seeds, or over five sampled chains of one run for GLoC-EHR (Reasoning); rows without SD are single runs or deterministic (LR). Best in bold, second underlined. ST: one model per task; MT: one model for all three tasks. C : code counts; T4096 : text embedding; ECLS / Emean : frozen EHR-BERT vectors; a slash names the classifier head.
AUROC ↑
Macro ↑
Model
Mortality
Long LOS
Readmission
AUROC
AUPRC
GLoC-EHR (MT)
0.9545±0.0013
0.8671±0.0030
0.7976±0.0052
0.8731±0.0030
0.5236±0.0088
without alignment
0.9470
0.8578
0.7817
0.8622
0.5354
Qwen3-1.7B + LoRA
0.9414
0.8411
0.7762
0.8529
0.4990
Appendix
Table 12: Single-run controls on the MIMIC-IV test set: AUROC per task and three-task macro AUROC/AUPRC. GLoC-EHR (MT) is the mean ± SD over five seeds from Table 1 . Without alignment: the same architecture trained from randomly initialized interface modules; Qwen3-1.7B + LoRA: the backbone fine-tuned on the serialized record. Both are single runs (Appendix C.1 ).
Model
Mortality
Long LOS
Readmission
Overall (%)
GLoC-EHR (Reasoning)
33,966
33,883
37,623
99.62
Qwen3-1.7B (CoT)
6,799
6,799
7,562
99.93
Qwen3-1.7B (CoT, our prompt)
6,380
6,529
7,139
94.68
Llama-3.3-70B (CoT)
6,804
6,804
7,566
100.00
Llama-3.3-70B (CoT, our prompt)
6,804
6,804
7,566
100.00
Appendix
Table 13: Answer-anchor coverage of generated chains (cases reaching a parseable answer; n = 6,804 / 6,804 / 7,566; GLoC-EHR (Reasoning) counts all five chains per case). For the zero-shot baselines, remaining cases are scored by forced closure.
Evaluation variant
n
Macro AUROC
Macro AUPRC
Single chains, valid anchors only (reported)
21,094
0.8656
0.4819
Single chains, no anchor ↦0.5
21,174
0.8654
0.4814
Average of five chains, valid anchors
21,161
0.8668
0.4897
Average of five chains, forced closure
21,174
0.8665
0.4897
Appendix
Table 14: Scoring sensitivities of GLoC-EHR (Reasoning); unweighted three-task macro. Single-chain rows are means over the five chains; n counts cases.
Support
Unsupported
Distinct
Macro
Model
Bullets
precision ↑
rate ↓
supported ↑
AUROC ↑
GLoC-EHR (Reasoning)
7.83
0.424±0.001
0.440±0.001
4.16±0.01
0.8656±0.0008
Qwen3-1.7B (CoT, our prompt)
6.69
0.452
0.420
3.88
0.5087
Llama-3.3-70B (CoT, our prompt)
6.85
0.752
0.060
6.33
0.5491
Appendix
Table 15: Evidence of GLoC-EHR (Reasoning) and the zero-shot “CoT, our prompt” baselines on the MIMIC-IV test set, scored with the reward matcher ( τ0=0.55 ) over chains that reach the answer anchor. GLoC-EHR (Reasoning): mean ± SD over five sampled chains, as in Table 2 . The zero-shot models read the serialized record; GLoC-EHR receives it only through its EHR memories. Macro AUROC is taken from Table 10 .
Figure 4: Evidence quality over GRPO training (training rollouts of all ranks, 200-step bins): (a) share of evidence bullets that match no event in the input window, (b) distinct supported findings per response, and (c) AUROC of the answer predicted from the bullets alone with the EHR memories switched off.
Next sodium < 135 mmol/L
Next platelet count < 150 K/ μ L
Method
AUROC
AUPRC
AUROC
AUPRC
GLoC-EHR (Reasoning)
0.657
0.500
0.629
0.499
Qwen3-1.7B, zero-shot
0.556
0.410
0.525
0.344
Appendix
Table 16: Zero-shot answers to two laboratory questions absent from training (held-out MIMIC-IV admissions; 750 positives and 1,500 length-matched negatives per question, so AUPRC is 0.33 by chance). GLoC-EHR (Reasoning) scores one sampled chain per case.
Patient portals now give individuals direct access to their electronic health records (EHRs), yet access alone does not ensure patients understand or act on the complex clinical information contained in these records. The ArchEHR-QA 2026 shared task addresses this challenge by focusing on grounded question answering over EHRs, and this paper presents the system developed by the HealthNLP_Retrievers team for this task. The proposed approach uses a multi-stage cascaded pipeline powered by the Gemini 2.5 Pro large language model to interpret patient-authored questions and retrieve relevant evidence from lengthy clinical notes. Our architecture comprises four integrated modules: (1) a few-shot query reformulation unit which summarizes verbose patient queries; (2) a heuristic-based evidence scorer which ranks clinical sentences to prioritize recall; (3) a grounded response generator which synthesizes professional-caliber answers restricted strictly to identified evidence; and (4) a high-precision many-to-many alignment framework which links generated answers to supporting clinical sentences. This cascaded approach achieved competitive results. Across the individual tracks, the system ranked 1st in question interpretation, 5th in answer generation, 7th in evidence identification, and 9th in answer-evidence alignment. These results show that integrating large language models within a structured multi-stage pipeline improves grounding, precision, and the professional quality of patient-oriented health communication. To support reproducibility, our source code is publicly available in our GitHub repository
Large language models (LLMs) exhibit strong natural-language reasoning abilities for clinical decision support, but struggle to effectively model structured longitudinal electronic health records (EHRs). In contrast, EHR foundation models can learn predictive patient representations, yet lack interpretable language-based reasoning. To bridge this gap, we propose ChatHealthAI, a multimodal reasoning framework that aligns structured EHR representations from a pretrained EHR foundation model with the semantic space of a frozen LLM through a task-aware resampler. By integrating longitudinal patient representations with refined clinical event descriptions, ChatHealthAI enables clinically grounded natural-language reasoning while maintaining accurate patient prediction. We evaluated ChatHealthAI on three clinical predictive tasks from the EHRSHOT benchmark. Results show that ChatHealthAI improves reasoning quality and interpretability while preserving competitive predictive performance. These findings highlight the potential of integrating EHR foundation models with pretrained LLMs for interpretable clinical prediction.
Bo-Hong Wang, Baicheng Peng, Ruilin Wang +3
School of Computer Science, McGill University, Montreal, QC, Canada · Mila - Quebec AI Institute, Montreal, QC, Canada
Longitudinal electronic health records (EHRs) capture years of patient history across notes, codes, labs, and procedures, and contain evidence needed to reason about likely clinical outcomes. However, comprehensive clinician review of these records is impractical, and LLM-based processing is costly and often unreliable, missing some relevant observations while hallucinating others. We therefore propose EviGen, a three-layer framework for verifiable clinical rationale generation that addresses these challenges. The first layer is a patient-conditioned retriever that uses learnable queries to find evidence predictive of, not just textually relevant to, a clinical outcome and ranks it by prediction attribution scores. The second layer is an LLM generator that consumes this ranked evidence as a scaffold to produce a clinical rationale grounded in the retrieved spans. The third layer is a process-supervised verifier that checks the generated rationale at the reasoning-step level, flagging unreliable claims. Across three medical prediction datasets, EviGen improves prediction performance and rationale faithfulness over full-context LLM and RAG baselines, and is preferred by clinical reviewers in a usability evaluation.