Frontier language models are rarely used in clinical workflows because the realistic, longitudinal benchmarks needed to develop them are scarce. Real electronic health record (EHR) data cannot be openly shared due to privacy, ethics or data use issues and it does not contain verifiable ground truth since the chart records only reflect what clinicians documented. We introduce Synthetic Hospital, an open, fully synthetic, fact-grounded longitudinal EHR benchmark that resolves the open sharing and verifiable ground truth barriers. Built entirely from public medical-education material with no protected health information, it comprises 1,268 longitudinal patients and 5,602 encounters, where every diagnosis, finding, and temporal relation is grounded in standard ontologies (ICD-10-CM, SNOMED CT, LOINC) and with a complete provenance chain back to its source medical education material. Synthetic Hospital is served through a simulated hospital record system that mirrors real EHR infrastructure (standard interoperability APIs, role-based access and function-calling interface). In a blinded review, physicians distinguished its records from real patient charts at near-chance rates (53%). Across 10 frontier and open models, none approaches ceiling: the best model reconstructs a patient's longitudinal problem list with a severity-weighted F1 of 0.73, level with the mean of seven physicians on a matched subset but well below the best of them (0.89), and misses roughly half of clinically relevant findings when summarizing a chart. Overall, these results highlight that Synthetic Hospital is a difficult and realistic test of clinical AI performance.
Figures & tables
Task
Input
Output
Primary Metric
Instances (public / held-out / train)
Patient diagnosis
Longitudinal EHR (multi-encounter)
Longitudinal problem list (ICD-10 + acuity)
Severity-weighted F1
200 / 268 / 800
Summarization
Clinical question + EHR sections
Structured clinical summary
Finding-level F1
200 / 268 / 800
Specialty summarization
EHR + target specialty
Specialty-focused summary
Specialty-relevance F1
983 / 1,325 / 4,037
Evidence retrieval
Diagnosis + patient record
Ranked evidence passages
Precision@5, NDCG@10
200 / 268 / 800
Imaging indication
Imaging order with vague indication + EHR
Inferred clinical question + pre-read
Question concept F1
276 / 407 / 1,182
Table 1: Synthetic Hospital benchmark tasks. For each task, we report the standardized input, expected output, primary evaluation metric, and the number of instances in the public, held-out, and training splits (12,014 in total). Secondary evaluation metrics are provided in Appendix A .
Patient Dx
Summ.
Spec. Summ.
Retrieval
Imaging
Overall
Model
Sev.-wgt. F1
Find. F1
Rel. F1
P@5
NDCG@10
Concept F1
Mean rank ↓
Gemini 3.1
0.681
0.502
0.582
0.809
0.497
0.496
4.40
GPT 5.3
0.703
0.489
0.595
0.833
0.536
0.518
2.40
Kimi 2.5-thinking (1T-A32B)
0.732
0.532
0.590
0.811
0.524
0.517
2.80
Opus 4.6
0.615
0.550
0.680
0.816
0.521
0.473
4.00
DeepSeek V3.2 (671B-A37B)
0.660
0.399
0.523
0.751
0.476
0.492
7.40
Table 2: Single-turn results across 10 models and five task variants using locked prompting strategies (public split). Bold = best; underline = second-best; higher is better except mean rank (lower is better). Summ. = whole-patient summarization; Spec. Summ. = specialty-conditioned summarization. Strategies: zero-shot (retrieval, Spec. Summ.); CoT (patient diagnosis); ontology-grounded (Summ.); few-shot (imaging). Mean rank averages ranks across the five tasks. Imaging is scored by concept F1 over graph-linked diagnoses and findings. ‡ Physician mean for seven physicians on a 13-patient subset; physicians did not complete Spec. Summ. Secondary metrics are in Appendix Table A .
Condition
Effect
Task (metric)
Model
LLM,
Agent,
Agent,
Agentic
Self-
full ctx.
full ctx.
self-retr.
loop
retrieval
Patient diagnosis
GPT 5.3
0.705
0.661
0.841
− 0.043
+ 0.180
(sev.-wgt. F1, chart-neutral)
Mistral Large
0.669
0.628
0.810
− 0.042
+ 0.183
Llama 4 Scout
0.554
0.794
0.894
+ 0.240
+ 0.099
Summarization
GPT 5.3
0.342
0.196
0.142
− 0.146
− 0.054
Table 3: Decomposition of agentic performance into loop and self-retrieval effects. Full-context LLM uses single-turn inference; full-context agent uses the same context within an agent loop; self-retrieving agent gathers evidence through the 13-tool EHR API. Loop effect = full-context agent − LLM; self-retrieval effect = self-retrieving − full-context agent. Positive values favor the more agentic condition; bold = best condition. Missing responses are scored zero.
Appendix figures & tables9 assets
Supplementary material from the paper’s appendix.
Appendix
Patient Dx
Summ.
Spec. Summ.
Retrieval
Imaging
Model
ICD Spec
Acuity
Prec
Rec
Omiss.
Hall.
Omiss.
Leak.
Abst.
Hall.
MAP@10
MRR
DiffCov
FindRec
Gemini 3.1
0.931
0.785
0.606
0.787
0.498
0.002
0.611
0.121
0.705
0.006
0.559
0.845
0.552
0.164
GPT 5.3
0.931
0.790
0.645
0.781
0.511
0.004
0.590
0.131
0.403
0.016
0.588
0.906
0.587
0.147
Kimi 2.5-thinking (1T-A32B)
0.940
0.657
0.687
0.787
0.468
0.001
0.580
0.143
0.682
0.002
0.575
0.887
0.583
0.124
Opus 4.6
0.934
0.696
0.535
0.738
0.450
0.001
0.467
0.202
0.708
0.038
0.575
0.895
0.574
0.142
DeepSeek V3.2 (671B-A37B)
0.924
0.798
0.653
0.682
0.601
0.001
0.647
0.119
0.710
0.014
0.529
0.811
0.575
0.140
Appendix
Appendix Table A: Secondary metrics for all 10 models (locked-strategy single-turn, public split). Includes ICD-10 specificity and chart-neutral precision (patient diagnosis) and summarization hallucination ( Summ. and Spec. Summ. Hall.), demoted here from primary Table 2 . Summ. Omiss. is whole-patient summarization omission under the structured strategy; the Spec. Summ. group reports specialty-conditioned omission, leakage (off-specialty inclusion), abstention accuracy on absent specialties, and hallucination. Bold = best per column; lower is better for omission, leakage, and hallucination, higher for all others. Hallucination columns (near-zero throughout) are not bolded.
Task
Model
Zero-shot
Structured
CoT
Few-shot
Patient Dx (Severity-weighted F1 score)
GPT 5.3
0.676
0.665
0.710
0.732
DeepSeek V3.2
0.664
0.633
0.633
0.652
Mistral Large
0.578
0.587
0.699
0.611
Summarization (Finding-level F1 score)
GPT 5.3
0.320
0.411
0.269
0.296
DeepSeek V3.2
0.215
0.323
0.175
0.209
Mistral Large
0.301
0.314
0.246
0.209
Appendix
Appendix Table C.1: Prompting-strategy ablation: 3 models × 4 strategies × 4 tasks (75-item pilot). Cells show the task’s primary metric (F1 for patient diagnosis; F1 for summarization; P@5 for retrieval; F1 for imaging). We see that no prompting strategy is best for any task. We also see that the right prompting strategy can make overall weaker models score better than frontier models. This means that controlling for prompting is important.
Δ = GPT-rendered − Kimi-rendered
Model
Pt. Dx
Summ.
Retr.
Img.
Mean ∣Δ∣
(Sev. F1, chart-neutral)
(Find. F1)
(P@5)
(concept F1)
Kimi 2.5-thinking (1T-A32B)
+0.025
−0.004
−0.074
−0.032
0.034
GPT 5.3
−0.010
−0.003
−0.144
−0.012
0.042
Opus 4.6
+0.003
+0.003
−0.087†
−0.020
0.028
GLM 5 (744B-A40B)
−0.021
+0.027
−0.058
−0.041
0.037
Appendix
Appendix Table C.2: Generator-robustness deltas on the 100-patient holdout. Each cell is the item-level paired change in the task’s holdout metric when the same clinical content is re-rendered by GPT 5.3 instead of Kimi k2.5 (negative = lower under GPT rendering); N=48 – 135 scored items per cell. † : Holm-corrected Wilcoxon p<0.05 . Kimi 2.5-thinking changes by less than 0.05 on three of four tasks, as do the generator-neutral anchors (Opus, GLM); on retrieval every model scores lower under GPT rendering ( −0.058 to −0.144 ), and GPT 5.3 itself drops the most, i.e. GPT does worse on its own rendering, the opposite of a home-field advantage; only the Opus retrieval change is Holm-significant. Patient diagnosis and summarization are stable under their primary metrics: all deltas fall within ±0.027 and none is Holm-significant. Bottom row: Spearman rank correlation of the four-model ordering across conditions; the ordering is preserved ( ρ≥0.80 ) on three of four tasks, with retrieval ( ρ=−0.40 ) the exception, reflecting reordering among models separated by small P@5 margins on the coarsest (rank-5) metric with the smallest per-cell N . The specialty-relevance task is omitted; it is not patient-scoped in this holdout.
Layer
Element
Count
Grounded
Nodes
Diagnoses
9,623
ICD-10-CM 75.9%; SNOMED 88.0%
Clinical findings
36,620
SNOMED 97.5%
of which lab values
3,488
LOINC 28.9%
Fact cards
53,999
85.6% linked
EHR sections
51,940
6,990 questions
Edges
Question–diagnosis
51,075
deterministic
Appendix
Appendix Table D: Knowledge graph size and ontology coverage after stage (2). ICD-10-CM coverage counts only table-validated codes; a further 14.6% of diagnoses carry a flagged, unvalidated LLM-proposed code.
Comparison
JSD
Spearman ρ
Synthetic Hospital vs. source corpus
0.029
0.83
Synthetic Hospital vs. Synthea
0.326
0.70
Source corpus vs. Synthea
0.436
–
Appendix
Appendix Table H.1: Distributional comparison of ICD-10 chapter frequencies across Synthetic Hospital, its source corpus, and Synthea. Synthetic Hospital closely preserves the case mix of its source material while differing substantially from the population-oriented Synthea cohort. Lower JSD indicates more similar distributions.
Figure 1: ICD-10 chapter distributions across Synthetic Hospital, its source corpus, and Synthea. Synthetic Hospital closely follows the education-derived case mix of its source corpus, whereas the population-oriented Synthea cohort is dominated by Z00–Z99 encounters.
ICD-10 chapter
Synth. Hospital (%)
Source corpus (%)
Synthea (%)
E00-E89 Endocrine/Metabolic
13.8
10.1
3.1
I00-I99 Circulatory
12.6
9.6
1.4
Z00-Z99 Factors/Health status
9.3
2.9
56.8
F01-F99 Mental/Behavioral
6.0
6.7
9.4
N00-N99 Genitourinary
5.9
6.0
1.3
A00-B99 Infectious
5.5
7.5
0.1
Appendix
Appendix Table H.2: ICD-10 chapter distributions underlying the case-mix comparison in Table H.1 .
Comorbidity pair
Odds ratio
T2DM – CKD (E11–N18)
14.3
AFib – Ischemic stroke (I48–I63)
16.01
HTN – Heart failure (I10–I50)
9.43
Hyperlipidemia – CAD (E78–I25)
21.63
COPD – Respiratory failure (J44–J96)
70.41
T2DM – CAD (E11–I25)
5.16
Appendix
Appendix Table H.3: Patient-level associations for ten canonical comorbidity pairs in Synthetic Hospital. Odds ratios use Haldane–Anscombe correction and 3-character ICD-10 categories. All ten pre-specified pairs have odds ratios above 4. Synthea comparisons are omitted because two pairs contain no mapped patients in either category, producing degenerate corrected estimates.
Benchmark
Paradigm
No Real Data
Provenance
Ground truth
Narr.
Synthea ( Walonoski et al., 2018 )
Rule-based
Yes
Module logic
API responses
No
EHR-Safe ( Yoon et al., 2023 )
GAN / statistical
No
None
Statistical
No
HALO ( Theodorou et al., 2023 ) CEHR-GPT ( Pang et al., 2024 )
Autoregressive
No
None
Distributional
No
SimSUM ( Rabaey et al., 2024 )
LLM-generated
Yes
Partial
Annotations
Yes
Zhou et al. ( Zhou et al., 2026 )
Knowledge-grounded
No
Partial
Statistical
No
Synthetic Hospital (ours)
Knowledge graph
Yes
Full
Ontology-grounded
Yes
Appendix
Appendix Table I: Synthetic clinical-data generation paradigms. All narrative approaches (SimSUM, ours) use an LLM to produce prose; the distinction is provenance of the ground truth. SimSUM annotates its own generated text, so generation errors can enter the labels; our ground truth is derived from the ontology-grounded graph independently of the generated narrative, so narrative errors cannot corrupt it. “No Real Data” marks approaches that can be built without any access to real patient records (Yes) or that require them (No). The Synthetic Hospital is the only approach that combines full provenance, precise ground truth, rich clinical narratives and fully synthetic data.
The generation of high-fidelity synthetic Electronic Health Records (EHR) is crucial for advancing medical research while preserving patient privacy. However, head-to-head comparison of existing generative models is hindered by disjointed codebases, incompatible data loaders, conflicting library dependencies, and inconsistent evaluation protocols. To address these gaps, we introduce a lightweight, end-to-end benchmarking framework for reproducible synthetic EHR evaluation, organized as a unified pipeline spanning data ingestion, standardized model training, and architecture-agnostic evaluation. Our current implementation targets the generation of longitudinal ICD diagnosis codes -- the most commonly studied modality in this literature -- and is built on the community-maintained PyHealth library. We reimplement and unify strong baselines (MedGAN, CorGAN, PromptEHR, HALO) under full ICD-9 vocabulary granularity, and add a lightweight GPT-2 baseline from the general-purpose sequence-modeling literature. We contribute a rigorous, architecture-agnostic privacy-utility evaluation suite that applies identically to GAN- and transformer-based generators, and report bootstrapped confidence intervals across all metrics. We further analyze the poor long-tailed performance of existing models and discuss the extensibility of our framework beyond diagnosis codes. By lowering the engineering barrier to running, extending, and evaluating under a single pipeline, we introduce a starting point for community-driven reproducibility and benchmarking synthetic EHR models.
Large language model (LLM)-based clinical assistants are increasingly being integrated into electronic health record (EHR) systems, transforming how clinicians retrieve and synthesize information from patient records. Their safety and utility depend on rigorous evaluation, yet existing benchmarks are manually curated, costly to update, and rapidly become obsolete with evolving technological advancements. We present a scalable framework that automatically generates question--answer pairs from longitudinal EHR notes. Nineteen clinicians validate the benchmark generator, producing the Benchmark for Retrieving Information in EHRs (BRIE), a continuously maintainable evaluation dataset. Across nine LLMs and five inference strategies, state-of-the-art systems frequently omit clinically important information, particularly for questions requiring synthesis across multiple documents and encounters. Because the generator itself is validated, BRIE supports evaluations that static benchmarks cannot, including the generation of multiple answers that reflect variation in clinician reasoning for robust performance assessment and continuously refreshing benchmark content to guard against leakage. Our results demonstrate that scalable benchmark generation enables rigorous, up-to-date evaluation of clinical LLMs as they are deployed in rapidly evolving healthcare settings.
Jordan L. Cahoon, Chloe O. Stanwyck, Sulaiman Somani +23
Access to clinical data is essential for developing reliable healthcare machine learning systems, but direct use of electronic health records is constrained by privacy regulation, institutional review, data-use agreements, and the risk of re-identification. Synthetic data promises a practical alternative: it can preserve useful statistical and clinical structure while reducing exposure of sensitive patient records. Prior studies often evaluate a single generator, one dataset, or a narrow downstream task, making it difficult to know when synthetic data can support model development and when it fails to preserve task-critical signal. We introduce CoMedBench, a reproducible benchmark that evaluates a family of generators under a common clinical-validity framework and one shared training and evaluation engine, spanning static tabular and temporal downstream tasks on established critical-care datasets. In total the benchmark spans 37 dataset-task pairs across two modalities consists of 20 static tabular and 17 temporal ICU time-series-drawn from seven public data sources: three intensive-care databases (MIMIC-III, MIMIC-IV, and eICU) together with the UCI Machine Learning Repository, the CDC BRFSS diabetes cohort (2015), NHANES (1999-2014), and the pycox survival datasets (GBSG and METABRIC). The benchmark evaluates both statistical fidelity and task utility by comparing models trained and tested across real and synthetic data. In these settings, synthetic training data preserves most of the downstream signal: on tabular tasks the reference generator CoMed-CTGAN retains a mean AUROC utility (the synthetic-to-real performance ratio) of 90.6%, rising to 97.3% for the strongest generator, CoMed-TVAE. Temporal ICU tasks are harder and more generator-sensitive: CoMed-CTGAN retains 81.6% (AUROC) and only 64.0% under the imbalance-sensitive AUPRC, whereas CoMed-TVAE still retains ~95% (AUROC).
Akanta Das, Al Amin Farhad, Mrinmoy Sarkar Anto +3