Bridging the EHR Divide: Asymmetric Contrastive Learning for Cross-National Medical Representation Transfer
Organizations: Information Networking Institute Carnegie Mellon University Pittsburgh, United States
Abstract
Cross-system transfer of longitudinal Electronic Health Record (EHR) representations is challenging because clinical coding, patient populations, and healthcare workflows differ substantially across institutions and countries. We introduce Asymmetric Supervised Contrastive Learning (Asymmetric SupCon), a task-specific pre-training objective motivated by the heterogeneity of negative clinical outcomes. The objective clusters patients sharing a target positive outcome without explicitly attracting negative trajectories toward one another. We pre-train temporal Transformer encoders on longitudinal records from 3.98 million patients in the Taiwanese National Health Insurance Research Database (NHIRD) and transfer them to two U.S. EHR datasets, MIMIC-IV and EHRSHOT. A hybrid semantic mapping pipeline combining direct mappings with embedding-based retrieval enables transfer across heterogeneous clinical vocabularies. On MIMIC-IV, NHIRD pre-training consistently improves over random initialization while substantially narrowing the performance gap to task-specific in-domain pre-training. On EHRSHOT, the transferred models show particularly strong few-shot performance for incident disease prediction. A controlled objective ablation under a matched pre-training scale shows that Asymmetric SupCon achieves higher mean AUPRC than direct supervised BCE transfer on all four evaluated tasks and Standard SupCon on three of four, with a 0.003 AUPRC deficit on readmission. These results support asymmetric contrastive pre-training as an effective approach for task-specific cross-national EHR representation transfer. Code is available at https://github.com/qingYzhang/Asymmetric_SupCon.
Figures & tables
| Task | Metric | Train-from-Scratch | In-Domain (MIMIC MIMIC) | Cross-Domain (NHIRD MIMIC) |
|---|---|---|---|---|
| 30-Day Readmission | AUROC | |||
| AUPRC | ||||
| F1 | ||||
| 90-Day Mortality | AUROC | |||
| AUPRC | ||||
| F1 |
| Task | Model | All | |||||
|---|---|---|---|---|---|---|---|
| Long LOS | Best Baseline | 0.576 | 0.282 | 0.315 | 0.337 | 0.439 | 0.473 |
| AsymSupCon | 0.517 | 0.298 | 0.333 | 0.355 | 0.431 | 0.450 | |
| ICU Admission | Best Baseline | 0.324 | 0.093 | 0.142 | 0.181 | 0.209 | 0.256 |
| AsymSupCon | 0.138 | 0.055 | 0.080 | 0.090 | 0.108 | 0.129 | |
| 30-Day Readmission | Best Baseline | 0.412 | 0.150 | 0.237 | 0.297 | 0.368 | 0.350 |
| AsymSupCon | 0.343 | 0.172 | 0.238 | 0.243 | 0.278 | 0.289 |
| MIMIC-IV | EHRSHOT ( ) | |||
|---|---|---|---|---|
| Pre-training Objective | 30-Day Readmission | 90-Day Mortality | Acute MI | Hyperlipidemia |
| Random Initialization | ||||
| Supervised BCE Transfer | ||||
| Masked Language Modeling (MLM) | ||||
| Standard SupCon | ||||
| Asymmetric SupCon (Ours) | ||||
Appendix figures & tables2 assets
Supplementary material from the paper’s appendix.
Appendix
| Threshold | Model | SNOMED | HCPCS | RxNorm | CPT-4 | LOINC | ICD-10-PCS | ICD-O-3 |
|---|---|---|---|---|---|---|---|---|
| Qwen3-Embedding-8B | 62.64% (4910/7838) | 88.10% (37/42) | 50.85% (1137/2236) | 27.19% (143/526) | 33.75% (243/720) | 70.59% (24/34) | 95.24% (20/21) | |
| SFR-Embedding-Mistral | 60.18% (4789/7958) | 82.00% (41/50) | 55.95% (1449/2590) | 30.50% (212/695) | 35.67% (305/855) | 71.43% (25/35) | 90.00% (18/20) | |
| Linq-Embed-Mistral | 66.32% (4780/7207) | 93.33% (14/15) | 62.44% (1350/2162) | 26.85% (109/406) | 41.68% (228/547) | 68.75% (11/16) | 95.00% (19/20) | |
| multilingual-e5-large-instruct | 57.69% (4386/7603) | 87.27% (48/55) | 57.61% (1446/2510) | 29.69% (198/667) | 42.45% (346/815) | 75.00% (27/36) | 90.00% (18/20) | |
| Qwen3-Embedding-8B † | 72.12% (4661/6463) | 90.91% (10/11) | 54.13% (1010/1866) | 41.62% (82/197) | 32.39% (161/497) | 71.43% (15/21) | 95.24% (20/21) | |
| SFR-Embedding-Mistral | 61.82% (4766/7709) | 90.00% (27/30) | 57.40% (1287/2242) | 27.69% (167/603) | 33.33% (250/750) | 67.74% (21/31) | 90.00% (18/20) |
| Task | Model | All | 1 | 2 | 4 | 8 | 12 | 16 | 24 | 32 | 48 | 64 | 128 |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Long LOS | CLMBR | 0.576 | 0.282 | 0.304 | 0.309 | 0.315 | 0.335 | 0.337 | 0.358 | 0.364 | 0.410 | 0.439 | 0.473 |
| GBM | 0.532 | 0.255 | 0.253 | 0.265 | 0.295 | 0.292 | 0.296 | 0.305 | 0.323 | 0.349 | 0.380 | 0.403 | |
| Logistic Regression | 0.384 | 0.237 | 0.251 | 0.283 | 0.283 | 0.283 | 0.294 | 0.308 | 0.311 | 0.346 | 0.340 | 0.365 | |
| Random Forest | 0.483 | 0.249 | 0.250 | 0.279 | 0.284 | 0.312 | 0.324 | 0.311 | 0.334 | 0.347 | 0.383 | 0.424 | |
| AsymSupCon (Ours) | 0.517 | 0.298 | 0.306 | 0.295 | 0.333 | 0.337 | 0.355 | 0.387 | 0.399 | 0.415 | 0.431 | 0.450 | |
| ICU Admission | CLMBR | 0.324 | 0.093 | 0.095 | 0.108 | 0.142 | 0.152 | 0.181 | 0.195 | 0.215 | 0.211 | 0.209 | 0.256 |