EHRAdapt: Adapting Pretrained Language Models to Electronic Health Records with Semantic Priors for Rare Clinical Events
Organizations: Lawrence Livermore National Laboratory · Kaiser Permanente
Abstract
Electronic health records (EHRs) encode clinical histories as (time, modality, code) tuples, whereas pretrained language models expect text tokens. Serializing them as text inflates sequence length and redundantly encodes structure. We introduce EHRAdapt, an adapter that maps tuples directly into a frozen language model's embedding space. Modality receives a learned embedding, time gaps enter through learned attention biases, and event codes receive dedicated vectors. Learning event vectors is the central challenge: clinical vocabularies are long-tailed, leaving rare events too few observations for reliable estimates. EHRAdapt therefore represents each event vector as the sum of a semantic prior and an evidence residual. The prior is a frozen embedding of the event's clinical description from a biomedical language model trained on clinical ontologies, mapped into the model's input space by a shared learned projection, so it supplies clinical meaning even when observations are scarce. The residual, a learned low-rank event-specific correction, refines it as evidence accumulates. We run continued pretraining on about 4 million patients' records with three frozen LLM backbones (OLMo2 1B, Llama3.2 1B, and OLMo2 7B), training only the adapter (0.1--0.6% of all parameters). The full adapter outperforms all ablations in held-out next-event prediction on every backbone. Removing the semantic pathway hurts rare events over ten times more than the most frequent ones, whereas removing the residual hurts overall prediction but improves it for the rarest events. On reportable infectious-disease and syndromic downstream classification tasks, EHRAdapt outperforms text-based LLM and count-based baselines, and both pathways improve rare-disease discrimination. The two pathways therefore play complementary roles, visible only when results are broken down by event frequency rather than averaged.
Figures & tables
| Configuration | OLMo2-1B | Llama3.2-1B | OLMo2-7B |
|---|---|---|---|
| Full model | |||
| Semantic prior only | |||
| Residual embeddings only | |||
| Frozen isometry + residual | |||
| Frozen isometry only |
| Model | All 14 | Diseases 13 | Frequent 7 | Rare 6 |
|---|---|---|---|---|
| BioLORD (mean) | ||||
| BioLORD (decay) | ||||
| OLMo 1B, native text | ||||
| Llama 1B, native text | ||||
| OLMo 7B, native text | ||||
| RF, tuned (count-based) |
| Configuration | CE | Freq. 7 | Rare 6 | GI/Resp. |
|---|---|---|---|---|
| OLMo 1B | ||||
| Semantic prior only | ||||
| Residual embeddings only | ||||
| Frozen isometry + residual | ||||
| Frozen isometry only | ||||
| Llama 1B | ||||
Appendix figures & tables12 assets
Supplementary material from the paper’s appendix.
Appendix
| Split | Fraction | Patients | Chunks | Clinical events |
|---|---|---|---|---|
| Train | 98% | 3.95M | 4.46M | 717M |
| Validation | 1% | 40,197 | 45,484 | 7.32M |
| Test | 1% | 39,734 | 45,144 | 7.32M |
| Backbone | Checkpoint | Layers | Heads | KV heads | Emb. norm | |
|---|---|---|---|---|---|---|
| OLMo2-1B | allenai/OLMo-2-0425-1B | 2,048 | 16 | 16 | 16 | 10.724 |
| Llama3.2-1B | meta-llama/Llama-3.2-1B | 2,048 | 16 | 32 | 8 | 0.987 |
| OLMo2-7B | allenai/OLMo-2-1124-7B | 4,096 | 32 | 32 | 32 | 7.423 |
| Size | OLMo2-1B | Llama3.2-1B | OLMo2-7B | |
|---|---|---|---|---|
| A. Totals | ||||
| Original text LM, including LM head | 1,484,916,736 | 1,235,814,400 | 7,298,617,344 | |
| Frozen backbone used by EHRAdapt | 1,279,395,840 | 1,235,814,400 | 6,887,575,552 | |
| Trainable EHRAdapt parameters | 6,821,685 | 6,825,781 | 8,556,341 | |
| Complete EHRAdapt model | 1,286,217,525 | 1,242,640,181 | 6,896,131,893 | |
| Trainable fraction | 0.53% | 0.55% | 0.12% | |
| Setting | OLMo2 1B | Llama3.2 1B | OLMo2 7B |
|---|---|---|---|
| Backbone optimization | Frozen | ||
| Optimizer | AdamW | ||
| Precision; attention | BF16; sdpa | ||
| Effective batch | 128 chunks (no gradient accumulation) | ||
| Learning rate: output/temporal | for vocabulary bias, logit scale, and temporal-attention bias | ||
| Learning rate: event encoder | for semantic projection, residual parameters, and modality embeddings | ||
| Class | 2019 | 2020 | 2021 | 2022 | Total |
|---|---|---|---|---|---|
| Gastrointestinal | 1,420 | 742 | 808 | 1,150 | 4,120 |
| Respiratory | 9,474 | 6,943 | 3,765 | 6,899 | 27,081 |
| Known | 10,894 | 7,685 | 4,573 | 8,049 | 31,201 |
| Control | 19,447 | 13,524 | 8,261 | 14,376 | 55,608 |
| All positive cases | 21,788 | 15,370 | 9,146 | 16,098 | 62,402 |
| Total | 41,235 | 28,894 | 17,407 | 30,474 | 118,010 |
| Class | 2019 | 2020 | 2021 | 2022 | Total |
| Campylobacteriosis | 825 | 448 | 479 | 634 | 2,386 |
| Salmonellosis (nontyphoidal) | 590 | 357 | 298 | 481 | 1,726 |
| Chickenpox (varicella) | 386 | 255 | 178 | 218 | 1,037 |
| Shigellosis | 419 | 180 | 246 | 360 | 1,205 |
| Pertussis | 486 | 97 | 28 | 30 | 641 |
| Invasive H. influenzae | 190 | 118 | 63 | 107 | 478 |
| Test 2020 | Test 2021 | Test 2022 | ||||
|---|---|---|---|---|---|---|
| Task | Train | Test | Train | Test | Train | Test |
| Syndromic | 41,235 | 28,894 | 70,129 | 17,407 | 87,536 | 30,474 |
| Infectious | 7,192 | 3,850 | 11,042 | 3,523 | 14,565 | 4,564 |
| Class | ICD-9-CM | ICD-10-CM |
| Syndromic classification | ||
| Gastrointestinal | – | A02.9, A03.0–A03.3, A03.8–A03.9, A04.5, A05.1, A05.3, A05.5, A32.11, A32.12, A32.7, A32.89, A32.9, B96.21 |
| Respiratory | 460, 461.9, 465.8, 465.9, 466.0, 486, 487.0, 487.1, 487.8, 488.01, 488.02, 488.09, 488.11, 488.12, 488.81, 488.82, 488.89, 490 | J00, J01.90, J06.9, J09.X1–J09.X3, J09.X9, J10.00, J10.01, J10.08, J10.1, J10.2, J10.81, J10.89, J11.00, J11.08, J11.1, J11.2, J11.89, J12.89, J12.9, J18.1, J18.8, J18.9, J20.9, J40 |
| Known | Any case diagnosis code outside the two syndromic code lists | |
| Control | Matched patients (same demographics and geographic location) with none of the codes above | |
| Infectious-disease classification | ||
| Configuration | CE | Top-1 (%) | Top-10 (%) |
|---|---|---|---|
| OLMo2-1B | |||
| Full model | |||
| Semantic prior only | |||
| Residual embeddings only | |||
| Frozen isometry + residual | |||
| Frozen isometry only | |||
| Configuration | Frequent 7 | Rare 6 | GI/Resp. |
|---|---|---|---|
| OLMo2-1B | |||
| Semantic prior only | |||
| Residual embeddings only | |||
| Frozen isometry + residual | |||
| Frozen isometry only | |||
| Llama3.2-1B | |||
| Configuration | All 14 | Diseases 13 | Frequent 7 | Rare 6 |
|---|---|---|---|---|
| OLMo2-1B | ||||
| Full model | ||||
| Semantic prior only | ||||
| Residual embeddings only | ||||
| Frozen isometry + residual | ||||
| Frozen isometry only | ||||
| Configuration | All 4 | GI | Respiratory |
|---|---|---|---|
| OLMo2-1B | |||
| Full model | |||
| Semantic prior only | |||
| Residual embeddings only | |||
| Frozen isometry + residual | |||
| Frozen isometry only | |||