Synthetic Hospital: An Open, Verifiable, Physician-Validated Longitudinal EHR Benchmark
Organizations: Carnegie Mellon University
Abstract
Frontier language models are rarely used in clinical workflows because the realistic, longitudinal benchmarks needed to develop them are scarce. Real electronic health record (EHR) data cannot be openly shared due to privacy, ethics or data use issues and it does not contain verifiable ground truth since the chart records only reflect what clinicians documented. We introduce Synthetic Hospital, an open, fully synthetic, fact-grounded longitudinal EHR benchmark that resolves the open sharing and verifiable ground truth barriers. Built entirely from public medical-education material with no protected health information, it comprises 1,268 longitudinal patients and 5,602 encounters, where every diagnosis, finding, and temporal relation is grounded in standard ontologies (ICD-10-CM, SNOMED CT, LOINC) and with a complete provenance chain back to its source medical education material. Synthetic Hospital is served through a simulated hospital record system that mirrors real EHR infrastructure (standard interoperability APIs, role-based access and function-calling interface). In a blinded review, physicians distinguished its records from real patient charts at near-chance rates (53%). Across 10 frontier and open models, none approaches ceiling: the best model reconstructs a patient's longitudinal problem list with a severity-weighted F1 of 0.73, level with the mean of seven physicians on a matched subset but well below the best of them (0.89), and misses roughly half of clinically relevant findings when summarizing a chart. Overall, these results highlight that Synthetic Hospital is a difficult and realistic test of clinical AI performance.
Figures & tables
| Task | Input | Output | Primary Metric | Instances (public / held-out / train) |
| Patient diagnosis | Longitudinal EHR (multi-encounter) | Longitudinal problem list (ICD-10 + acuity) | Severity-weighted F1 | 200 / 268 / 800 |
| Summarization | Clinical question + EHR sections | Structured clinical summary | Finding-level F1 | 200 / 268 / 800 |
| Specialty summarization | EHR + target specialty | Specialty-focused summary | Specialty-relevance F1 | 983 / 1,325 / 4,037 |
| Evidence retrieval | Diagnosis + patient record | Ranked evidence passages | Precision@5, NDCG@10 | 200 / 268 / 800 |
| Imaging indication | Imaging order with vague indication + EHR | Inferred clinical question + pre-read | Question concept F1 | 276 / 407 / 1,182 |
| Patient Dx | Summ. | Spec. Summ. | Retrieval | Imaging | Overall | ||
| Model | Sev.-wgt. F1 | Find. F1 | Rel. F1 | P@5 | NDCG@10 | Concept F1 | Mean rank |
| Gemini 3.1 | 0.681 | 0.502 | 0.582 | 0.809 | 0.497 | 0.496 | 4.40 |
| GPT 5.3 | 0.703 | 0.489 | 0.595 | 0.833 | 0.536 | 0.518 | 2.40 |
| Kimi 2.5-thinking (1T-A32B) | 0.732 | 0.532 | 0.590 | 0.811 | 0.524 | 0.517 | 2.80 |
| Opus 4.6 | 0.615 | 0.550 | 0.680 | 0.816 | 0.521 | 0.473 | 4.00 |
| DeepSeek V3.2 (671B-A37B) | 0.660 | 0.399 | 0.523 | 0.751 | 0.476 | 0.492 | 7.40 |
| Condition | Effect | |||||
| Task (metric) | Model | LLM, | Agent, | Agent, | Agentic | Self- |
| full ctx. | full ctx. | self-retr. | loop | retrieval | ||
| Patient diagnosis | GPT 5.3 | 0.705 | 0.661 | 0.841 | 0.043 | 0.180 |
| (sev.-wgt. F1, chart-neutral) | Mistral Large | 0.669 | 0.628 | 0.810 | 0.042 | 0.183 |
| Llama 4 Scout | 0.554 | 0.794 | 0.894 | 0.240 | 0.099 | |
| Summarization | GPT 5.3 | 0.342 | 0.196 | 0.142 | 0.146 | 0.054 |
Appendix figures & tables9 assets
Supplementary material from the paper’s appendix.
Appendix
| Patient Dx | Summ. | Spec. Summ. | Retrieval | Imaging | ||||||||||
| Model | ICD Spec | Acuity | Prec | Rec | Omiss. | Hall. | Omiss. | Leak. | Abst. | Hall. | MAP@10 | MRR | DiffCov | FindRec |
| Gemini 3.1 | 0.931 | 0.785 | 0.606 | 0.787 | 0.498 | 0.002 | 0.611 | 0.121 | 0.705 | 0.006 | 0.559 | 0.845 | 0.552 | 0.164 |
| GPT 5.3 | 0.931 | 0.790 | 0.645 | 0.781 | 0.511 | 0.004 | 0.590 | 0.131 | 0.403 | 0.016 | 0.588 | 0.906 | 0.587 | 0.147 |
| Kimi 2.5-thinking (1T-A32B) | 0.940 | 0.657 | 0.687 | 0.787 | 0.468 | 0.001 | 0.580 | 0.143 | 0.682 | 0.002 | 0.575 | 0.887 | 0.583 | 0.124 |
| Opus 4.6 | 0.934 | 0.696 | 0.535 | 0.738 | 0.450 | 0.001 | 0.467 | 0.202 | 0.708 | 0.038 | 0.575 | 0.895 | 0.574 | 0.142 |
| DeepSeek V3.2 (671B-A37B) | 0.924 | 0.798 | 0.653 | 0.682 | 0.601 | 0.001 | 0.647 | 0.119 | 0.710 | 0.014 | 0.529 | 0.811 | 0.575 | 0.140 |
| Task | Model | Zero-shot | Structured | CoT | Few-shot |
| Patient Dx (Severity-weighted F1 score) | GPT 5.3 | 0.676 | 0.665 | 0.710 | 0.732 |
| DeepSeek V3.2 | 0.664 | 0.633 | 0.633 | 0.652 | |
| Mistral Large | 0.578 | 0.587 | 0.699 | 0.611 | |
| Summarization (Finding-level F1 score) | GPT 5.3 | 0.320 | 0.411 | 0.269 | 0.296 |
| DeepSeek V3.2 | 0.215 | 0.323 | 0.175 | 0.209 | |
| Mistral Large | 0.301 | 0.314 | 0.246 | 0.209 |
| = GPT-rendered Kimi-rendered | |||||
| Model | Pt. Dx | Summ. | Retr. | Img. | Mean |
| (Sev. F1, chart-neutral) | (Find. F1) | (P@5) | (concept F1) | ||
| Kimi 2.5-thinking (1T-A32B) | 0.034 | ||||
| GPT 5.3 | 0.042 | ||||
| Opus 4.6 | 0.028 | ||||
| GLM 5 (744B-A40B) | 0.037 | ||||
| Layer | Element | Count | Grounded |
| Nodes | Diagnoses | 9,623 | ICD-10-CM 75.9%; SNOMED 88.0% |
| Clinical findings | 36,620 | SNOMED 97.5% | |
| of which lab values | 3,488 | LOINC 28.9% | |
| Fact cards | 53,999 | 85.6% linked | |
| EHR sections | 51,940 | 6,990 questions | |
| Edges | Question–diagnosis | 51,075 | deterministic |
| Comparison | JSD | Spearman |
| Synthetic Hospital vs. source corpus | 0.029 | 0.83 |
| Synthetic Hospital vs. Synthea | 0.326 | 0.70 |
| Source corpus vs. Synthea | 0.436 | – |
| ICD-10 chapter | Synth. Hospital (%) | Source corpus (%) | Synthea (%) |
| E00-E89 Endocrine/Metabolic | 13.8 | 10.1 | 3.1 |
| I00-I99 Circulatory | 12.6 | 9.6 | 1.4 |
| Z00-Z99 Factors/Health status | 9.3 | 2.9 | 56.8 |
| F01-F99 Mental/Behavioral | 6.0 | 6.7 | 9.4 |
| N00-N99 Genitourinary | 5.9 | 6.0 | 1.3 |
| A00-B99 Infectious | 5.5 | 7.5 | 0.1 |
| Comorbidity pair | Odds ratio |
| T2DM – CKD (E11–N18) | 14.3 |
| AFib – Ischemic stroke (I48–I63) | 16.01 |
| HTN – Heart failure (I10–I50) | 9.43 |
| Hyperlipidemia – CAD (E78–I25) | 21.63 |
| COPD – Respiratory failure (J44–J96) | 70.41 |
| T2DM – CAD (E11–I25) | 5.16 |
| Benchmark | Paradigm | No Real Data | Provenance | Ground truth | Narr. |
| Synthea ( Walonoski et al., 2018 ) | Rule-based | Yes | Module logic | API responses | No |
| EHR-Safe ( Yoon et al., 2023 ) | GAN / statistical | No | None | Statistical | No |
| HALO ( Theodorou et al., 2023 ) CEHR-GPT ( Pang et al., 2024 ) | Autoregressive | No | None | Distributional | No |
| SimSUM ( Rabaey et al., 2024 ) | LLM-generated | Yes | Partial | Annotations | Yes |
| Zhou et al. ( Zhou et al., 2026 ) | Knowledge-grounded | No | Partial | Statistical | No |
| Synthetic Hospital (ours) | Knowledge graph | Yes | Full | Ontology-grounded | Yes |