cs.CLJun 17, 2026

Are LLMs Ready to Assist Physicians? PhysAssistBench for Interactive Doctor-Patient-EHR Assistance

Authors: Tianming DuPeijie YuSihan ShangDanli ShiMy Linh NguyenShengbo GaoGuangyuan LiYinghong Yu+7 more

Organizations: 1ELLIS Institute Finland · 2Aalto University · 3Tencent · 4Harbin Institute of Technology, Shenzhen · 5Hong Kong Polytechnic University · University of Oulu · 7Polytechnic University of Milan · 8Aarhus University · University of Turku · 10Technical University of Munich

Abstract

The most plausible near-term role of medical LLMs is to assist rather than replace physicians, yet current evaluations often test isolated capabilities: clinical knowledge, EHR system interaction, or patient communication. Physician assistance instead requires coordinating these capabilities within the same interaction, where physicians issue underspecified requests, patients describe symptoms ambiguously, and EHR systems demand precise tool use. We introduce PhysAssistBench, a benchmark for interactive doctor-patient-EHR assistance. Built from real MIMIC-IV cases, PhysAssistBench uses a scalable pipeline to construct agentic patients: interactive, record-grounded agents that turn static EHR records into multi-turn clinical scenarios while preserving clinical factuality. PhysAssistBench provides a curated bilingual evaluation set of 1,296 manually reviewed and physician-validated turns. Experiments with leading LLMs show that current models remain unreliable in this setting, which exposes a key bottleneck for clinical LLMs: reliable assistance requires coordination across knowledge, communication, and systems, not isolated gains in any of them.

Explore similar work

CardsList