In hospital workflows, electronic health records (EHRs) are often noisy, and may not contain the evidence needed to confirm events or measurements referenced in a clinical query. Even when database retrieval succeeds, clinical agents can overlook such discrepancies and return plausible but unsupported answers. We introduce EHR-RobustGym, a scalable and interactive environment for evaluating and training robust clinical agents grounded in noisy EHRs. Built on MIMIC-IV hospital records (365K patients, 31 tables, and over 500M records), EHR-RobustGym comprises 5,486 Clean-Noise pairs spanning six clinical intents and both patient-level and population-level queries. The pairs test robustness to Record-level, Value-level, and Query-level noise, while interactive SQL/Python execution and outcome verification support trajectory collection and training. Evaluating multiple LLMs reveals substantial robustness gaps: average task success across proprietary and large-scale open-weight models drops from 62.2% on Clean questions to 37.9% on Noise questions. At k=4, pass^k consistency falls below 50% for most evaluated models, exposing instability in clinical task completion. Supervised fine-tuning and reinforcement learning in EHR-RobustGym improve performance, with gains generalizing to five external EHR benchmarks. Together, these results position EHR-RobustGym as a testbed for evaluating and improving the evidence-grounded robustness of clinical agents.
Figures & tables
Figure 1: Performance Comparisons under Clean and Noise Conditions. (Left): Base success rates (%) for selected models across six clinical intents. (Right): A Query-level Noise request asks for the first insulin dose administered after blood potassium first exceeds 6.0 mmol/L. The incorrect response reports 10 Units given at 09:00, overlooking that the qualifying potassium result occurs at 10:00. In contrast, the evidence-grounded response returns NULL because the EHR contains no insulin administration after that result.
Dataset
Agentic Mode
Robustness Evaluation
Database Scale
Task Instances
Noise Conditions
Trajectory-level Evaluation
Patients
Tables
Elements
Types
Test
Train
MedAgentBench ( Jiang et al., 2025 )
✔
✗
✔
100
-
700K
10
300
-
EHR-Complex ( Qiao et al., 2026 )
✔
✗
✔
365K
31
>500M
12
3,915
48,092
TREQS ( Wang et al., 2020 )
✗
✗
✗
100
5
2.5M
4
996
8,988
EHRSQL (eICU) ( Lee et al., 2022 )
✗
✗ †
✗
<1K
10
1.5M
9
611
6,213
EHRSQL (MIMIC-III) ( Lee et al., 2022 )
✗
✗ †
✗
<1K
17
1.4M
9
1,122
9,318
Table 1: Comparison with existing EHR benchmarks.
Figure 2: Construction pipeline and task composition of EHR-RobustGym . Step-1. Based on chart-review needs and hospital workflows, clinical requests are derived. Step-2. These requests are grounded in executable evidence contracts with reference SQL and verified answers. Step-3. Clean -Noise pairs are created by changing one necessary condition without modifying the database. Step-4. The paired questions are jointly rewritten with fixed contracts, then verified by SQL execution and semantic checks. The right panel shows clinical intents, noise types, and data splits.
Figure 3: Overview of EHR-RobustGym . It contains paired Clean and Noise tasks with interactive SQL/Python execution and outcome verification for evaluating and training LLM agents.
Model
Demographics
Vitals
Medications
Cost
Labs
Diagnoses
Avg.
Proprietary API models
GPT-5.4 (xhigh)
76.2 / 47.7
63.6 / 48.5
63.2 / 53.8
95.7 / 65.2
91.3 / 58.2
71.4 / 58.2
78.3 / 55.8
GPT-4.1
28.5 / 17.7
18.2 / 18.2
38.7 / 31.6
17.4 / 17.4
34.1 / 26.7
53.6 / 37.7
37.7 / 28.5
Gemini-3.1-Pro
89.2 / 41.5
72.7 / 30.3
60.4 / 48.1
95.7 / 60.9
93.8 / 68.5
84.1 / 64.5
83.4 / 58.4
Claude Sonnet 4.6
62.3 / 28.5
60.6 / 24.2
58.5 / 22.6
65.2 / 13.0
64.4 / 19.2
63.2 / 42.7
62.5 / 26.3
Open-weight models ≥ 100B
Table 2: Task success rates of unadapted EHR agents on EHR-RobustGym .
Model
Method
Demographics
Vitals
Medications
Cost
Labs
Diagnoses
Avg.
Δ vs. Base
Prompt-only mitigation (API-based models)
GPT-5.4 (xhigh)
PE
81.5 / 43.8
69.7 / 45.5
65.1 / 52.4
82.6 / 52.2
91.5 / 59.5
78.2 / 59.1
80.9 / 55.3
+2.6 / -0.5
GPT-4.1
PE
30.0 / 23.1
18.2 / 6.1
38.2 / 31.1
21.7 / 21.7
35.4 / 23.1
53.2 / 40.5
38.3 / 28.0
+0.6 / -0.5
Gemini-3.1-Pro
PE
89.2 / 37.7
75.8 / 48.5
63.2 / 56.1
91.3 / 65.2
95.6 / 70.8
85.0 / 66.4
84.9 / 61.6
+1.5 / +3.2
Claude Sonnet 4.6
PE
66.2 / 33.8
78.8 / 24.2
63.2 / 25.0
87.0 / 17.4
68.2 / 29.0
70.0 / 47.3
68.1 / 32.3
+5.6 / +6.1
Kimi-K2.5
PE
77.7 / 16.2
87.9 / 18.2
59.9 / 33.5
91.3 / 8.7
93.8 / 51.0
80.9 / 49.5
81.5 / 40.5
-0.9 / +3.4
Table 3: EHR agent performance after prompting and training with EHR-RobustGym .
Figure 4: Scaling and Self-Improvement in EHR-RobustGym .
Figure 5: General and Medical Capabilities after EHR Training.
Figure 6: External EHR Generalization and Training Dynamics. (a) Qwen3-14B success rates on five external EHR benchmarks. (b) SFT loss and (c) subsequent GRPO reward and response length.
Clinical agents promise to democratize access to electronic health records (EHRs), yet existing benchmarks fail to reflect the complexity of practical EHR analysis, e.g., often operating on idealized, clean EHRs via static SQL generation rather than interactive execution. In this work, we introduce EHR-Complex, a large-scale benchmark designed for interactive clinical database reasoning. Built on the large MIMIC-IV substrate (365K patients, 31 tables, 500M+ records), EHR-Complex comprises about 52K tasks spanning six clinical intents, supporting both patient-level and population-level queries, where each task requires an agent to interact with a sandboxed environment by executing SQL queries or Python code. Notably, EHR-Complex considers the real-world SQL task complexity for longitudinal multi-table aggregation and compositional reasoning, resulting in 31.93 SQL structural components per query on average. Evaluation results on EHR-Complex reveal the clinical difficulty of these EHR reasoning scenarios, with the top-performing model achieving only 62.3% exact-match accuracy. Pass^k consistency drops below 50% for nearly all evaluated models at k=4, exposing broad stochastic fragility. A fine-grained analysis of more than 3,800 failed trajectories for representative LLMs reveals three dominant failure modes: SQL logic errors, medical-code lookup failures, and semantic misunderstandings. EHR-Complex provides a rigorous testbed for clinical agents and highlights remaining gaps in robust reasoning for large-scale EHR analysis.
Yitong Qiao, Lei Liu, Yue Shen +4
1Zhejiang University · †Work done during an internship at Ant Group. · 2Ant Group
We introduce PhysicianBench, a benchmark for evaluating LLM agents on physician tasks grounded in real clinical setting within electronic health record (EHR) environments. Existing medical agent benchmarks primarily focus on static knowledge recall, single-step atomic actions, or action intent without verifiable execution against the environment. As a result, they fail to capture the long-horizon, composite workflows that characterize real clinical systems. PhysicianBench comprises 100 long-horizon tasks adapted from real consultation cases between primary care and subspecialty physicians, with each task independently reviewed by a separate panel of physicians. Tasks are instantiated in an EHR environment with real patient records and accessed through the same standard APIs used by commercial EHR vendors. Tasks span 21 specialties (e.g., cardiology, endocrinology, oncology, psychiatry) and diverse workflow types (e.g., diagnosis interpretation, medication prescribing, treatment planning), requiring an average of 27 tool calls per task. Solving each task requires retrieving data across encounters, reasoning over heterogeneous clinical information, executing consequential clinical actions, and producing clinical documentation. Each task is decomposed into structured checkpoints (670 in total across the benchmark) capturing distinct stages of completion graded by task-specific scripts with execution-grounded verification. Across 13 proprietary and open-source LLM agents, the best-performing model achieves only 46% success rate (pass@1), while open-source models reach at most 19%, revealing a substantial gap between current agent capabilities and the demands of real-world clinical workflows. PhysicianBench provides a realistic and execution-grounded benchmark for measuring progress toward autonomous clinical agents.
Ruoqi Liu, Imran Q. Mohiuddin, Austin J. Schoeffler +10
Longitudinal electronic health records (EHRs) capture years of patient history across notes, codes, labs, and procedures, and contain evidence needed to reason about likely clinical outcomes. However, comprehensive clinician review of these records is impractical, and LLM-based processing is costly and often unreliable, missing some relevant observations while hallucinating others. We therefore propose EviGen, a three-layer framework for verifiable clinical rationale generation that addresses these challenges. The first layer is a patient-conditioned retriever that uses learnable queries to find evidence predictive of, not just textually relevant to, a clinical outcome and ranks it by prediction attribution scores. The second layer is an LLM generator that consumes this ranked evidence as a scaffold to produce a clinical rationale grounded in the retrieved spans. The third layer is a process-supervised verifier that checks the generated rationale at the reasoning-step level, flagging unreliable claims. Across three medical prediction datasets, EviGen improves prediction performance and rationale faithfulness over full-context LLM and RAG baselines, and is preferred by clinical reviewers in a usability evaluation.