GLoC-EHR: Evidence-Cited Clinical Reasoning over Global Context and Local EHR Events
Organizations: Interdisciplinary Program of Medical Informatics, Seoul National University College of Medicine · Department of Transdisciplinary Medicine, Seoul National University Hospital · Center for Data Science, Healthcare AI Research Institute, Seoul National University Hospital · Department of Medicine, Seoul National University College of Medicine
Abstract
Structured electronic health records (EHRs) contain a patient's clinical trajectory as a sequence of clinical codes. Answering clinical questions from such records requires both the context of the whole trajectory and the specific events that support the answer. We introduce GLoC-EHR, a multimodal language model that reads a contextual encoding of the record through a fixed-size global memory of the trajectory and a local memory of selected events. The model learns to generate hospital-course summaries from the global memory and descriptions of masked concepts from the local memory, aligning both with clinical text. It is then trained to cite evidence before answering, through rationale fine-tuning followed by group relative policy optimization (GRPO) with rewards for correct answers and record-supported evidence. On three MIMIC-IV outcome tasks, GLoC-EHR attains the highest macro AUROC among the compared models when it answers directly, whereas zero-shot LLMs reading the serialized record fall far behind. With evidence-cited reasoning, it stays close to its direct multi-task counterpart in macro AUROC, and the evidence terms of the objective reduce unsupported evidence at a similar macro AUROC. The local memory adds distinct supported findings, particularly under strict matching, without a detectable change in macro AUROC. Without retraining, GLoC-EHR transfers to EHRSHOT on par with EHR-BERT and answers two unseen laboratory questions better than zero-shot prompting of its own backbone.
Figures & tables
| AUROC | Macro (3 tasks) | ||||
| Method | Mortality | Long LOS | Readmission | AUROC | AUPRC |
| (24 h) | (24 h) | (full) | |||
| Zero-shot LLMs on serialized records | |||||
| Qwen3-1.7B (direct) | |||||
| Qwen3-1.7B (CoT) | |||||
| Llama-3.3-70B (direct) | |||||
| Macro | Macro | Unsupported | Distinct supported | Bullets-only | ||
|---|---|---|---|---|---|---|
| Variant | AUROC | AUPRC | rate | AUROC | ||
| GLoC-EHR (Reasoning) | ||||||
| one chain | 0.8661 | 0.4810 | 0.440 | 4.15 | 2.73 | 0.755 |
| w/o evidence terms | 0.8665 | 0.4896 | 0.577 | 3.66 | 2.21 | 0.707 |
| global route only | 0.8641 | 0.4783 | 0.431 | 4.00 | 2.41 | 0.767 |
| rationale SFT only | 0.6578 | 0.2013 | 0.582 | 3.92 | 2.34 | 0.690 |
| AUROC | Macro | ||||
|---|---|---|---|---|---|
| Method | Mortality | Long LOS | Readmission | AUROC | AUPRC |
| EHR-BERT (MT) | |||||
| code IDs only | |||||
| GLoC-EHR (MT) | |||||
| GLoC-EHR (Reasoning) | |||||
Appendix figures & tables15 assets
Supplementary material from the paper’s appendix.
Appendix
| Isotropy | Qualified concepts | |||||||
|---|---|---|---|---|---|---|---|---|
| Representation | Post-processing | PC 1 % | NN c % | NN q % | ||||
| Input-token mean | none | 295 | 15.3 | 0.29 | 0.97 | 0.61 | 75.3 | 48.4 |
| Input-token mean | whiten | 1,793 | 0.2 | 0.00 | 0.94 | 0.00 | 91.1 | 36.4 |
| Final hidden state | none | 70 | 21.3 | 0.89 | — | — | — | — |
| Final hidden state | top-PC removal | 116 | 23.4 | 0.00 | 0.99 | 0.36 | 87.0 | 39.5 |
| Final hidden state | whiten (used) | 1,908 | 0.3 | 0.00 | 0.94 | 0.00 | 98.4 | 31.9 |
| Stage | Supervision | Trainable components |
|---|---|---|
| Semantic encoder | masked concepts | EHR encoder and input projections |
| Global alignment | BHC generation | resampler, global cross-attention |
| Local alignment | concept descriptions + BHC replay | projection, selector, local cross-attention; local cross-attention only after hard selection |
| Direct prediction | binary outcome | global and local cross-attention |
| Rationale SFT, soft | teacher rationale + answer | base LLM, resampler, global and local cross-attention, local projection, selector |
| Rationale SFT, hard | teacher rationale + answer | base LLM, resampler, global and local cross-attention |
| Validation target | Full | Local off | Global off | Both off | |
|---|---|---|---|---|---|
| Concept description | 256 | 0.2797 | 3.3105 ( 1083.5%) | 0.5378 ( 92.3%) | 3.6285 |
| 128 | 0.2844 | 0.5536 ( 94.7%) | |||
| 64 | 0.3010 | 0.5913 ( 96.5%) | |||
| BHC generation | 256 | 1.7937 | 1.7922 ( 0.1%) | 3.2044 ( 78.6%) | — |
| 64 | 1.7942 | — | — |
| Eligible events | Selection inactive ( events) | ||||
|---|---|---|---|---|---|
| Task | Median | p90 | |||
| Mortality, long LOS (24 h) | 96.5 | 603 | 78.9% | 66.5% | 29.6% |
| Readmission (full) | 194 | 2,048 | 58.5% | 37.8% | 20.3% |
| Mortality | Long LOS | Readmission | All | |
| Cases (positives) | 453 (17) | 441 (152) | 1,168 (75) | 2,062 (244) |
| Persistence, cited mask | 0.874 | 0.885 | 0.851 | 0.863 |
| Persistence, control mask | 0.885 | 0.889 | 0.875 | 0.880 |
| Persistence, cited control | [ , ] | [ , ] | [ , ] | [ , ] |
| , cited / control | 0.007 / 0.007 | 0.021 / 0.021 | 0.012 / 0.012 | 0.013 / 0.013 |
| Label flips (%), cited / control | 1.1 / 1.3 | 1.8 / 2.9 | 0.4 / 0.3 | 0.9 / 1.1 |
| MIMIC-IV | |||||
|---|---|---|---|---|---|
| Task | Input | Training | Validation | Test | EHRSHOT |
| In-hospital mortality | First 24 h | 24,876 (3.19%) | 3,549 (3.18%) | 6,804 (3.15%) | 9,922 (2.75%) |
| Long LOS ( 7 d) | First 24 h | 24,876 (37.06%) | 3,549 (37.70%) | 6,804 (36.67%) | 9,922 (31.14%) |
| 30-day readmission | Full admission | 27,420 (7.55%) | 3,890 (7.46%) | 7,566 (6.89%) | 12,342 (23.70%) |
| Method | Mortality (24 h) | Long LOS (24 h) | Readmission (full) | Macro (3 tasks) |
|---|---|---|---|---|
| Zero-shot LLMs on serialized records (single run) | ||||
| Qwen3-1.7B (direct) | ||||
| Qwen3-1.7B (CoT) | ||||
| Qwen3-1.7B (CoT, our prompt) | ||||
| Llama-3.3-70B (direct) | ||||
| Llama-3.3-70B (CoT) | ||||
| Method | Mortality (24 h) | Long LOS (24 h) | Readmission (full) | Macro (3 tasks) |
|---|---|---|---|---|
| Zero-shot LLMs on serialized records (single run) | ||||
| Qwen3-1.7B (direct) | ||||
| Qwen3-1.7B (CoT) | ||||
| Qwen3-1.7B (CoT, our prompt) | ||||
| Llama-3.3-70B (direct) | ||||
| Llama-3.3-70B (CoT) | ||||
| AUROC | Macro | ||||
| Model | Mortality | Long LOS | Readmission | AUROC | AUPRC |
| GLoC-EHR (MT) | |||||
| without alignment | 0.9470 | 0.8578 | 0.7817 | 0.8622 | 0.5354 |
| Qwen3-1.7B + LoRA | 0.9414 | 0.8411 | 0.7762 | 0.8529 | 0.4990 |
| Model | Mortality | Long LOS | Readmission | Overall (%) |
|---|---|---|---|---|
| GLoC-EHR (Reasoning) | 33,966 | 33,883 | 37,623 | 99.62 |
| Qwen3-1.7B (CoT) | 6,799 | 6,799 | 7,562 | 99.93 |
| Qwen3-1.7B (CoT, our prompt) | 6,380 | 6,529 | 7,139 | 94.68 |
| Llama-3.3-70B (CoT) | 6,804 | 6,804 | 7,566 | 100.00 |
| Llama-3.3-70B (CoT, our prompt) | 6,804 | 6,804 | 7,566 | 100.00 |
| Evaluation variant | Macro AUROC | Macro AUPRC | |
|---|---|---|---|
| Single chains, valid anchors only (reported) | 21,094 | 0.8656 | 0.4819 |
| Single chains, no anchor | 21,174 | 0.8654 | 0.4814 |
| Average of five chains, valid anchors | 21,161 | 0.8668 | 0.4897 |
| Average of five chains, forced closure | 21,174 | 0.8665 | 0.4897 |
| Support | Unsupported | Distinct | Macro | ||
| Model | Bullets | precision | rate | supported | AUROC |
| GLoC-EHR (Reasoning) | 7.83 | ||||
| Qwen3-1.7B (CoT, our prompt) | 6.69 | 0.452 | 0.420 | 3.88 | 0.5087 |
| Llama-3.3-70B (CoT, our prompt) | 6.85 | 0.752 | 0.060 | 6.33 | 0.5491 |
| Next sodium 135 mmol/L | Next platelet count 150 K/ L | |||
|---|---|---|---|---|
| Method | AUROC | AUPRC | AUROC | AUPRC |
| GLoC-EHR (Reasoning) | 0.657 | 0.500 | 0.629 | 0.499 |
| Qwen3-1.7B, zero-shot | 0.556 | 0.410 | 0.525 | 0.344 |