Organizations: Peking University, Beijing, China · Northeastern University, Boston, MA, USA · Sun Yat-sen University, Guangzhou, China · Department of Medical Bioinformatics, School of Basic Medical Sciences, Peking University, Beijing, China
Proactive medical dialogue requires an agent to decide what to ask from incomplete patient information. Existing information-seeking approaches commonly prioritize questions that most reduce diagnostic uncertainty, but this criterion overlooks an important property of medical diagnosis: different diagnostic errors can carry substantially different consequences. The most informative question may therefore differ from the one most valuable for the downstream decision. We propose Expected-Severity-Risk (ESR), a consequence-aware question-supervision objective that values each candidate by its expected reduction in severity-aware terminal risk. Because questions must be selected before their answers are observed, ESR marginalizes over possible answers using train-only population statistics. Its rankings are then distilled into a prefix-only language policy, requiring no teacher-side risk computation at deployment. Across three matched Qwen3-4B training seeds on DDxPlus, ESR reduces mean high-severity diagnostic miss from 0.0645 to 0.0455 (29.5% relative reduction) and improves mean diagnostic accuracy from 0.9123 to 0.9320 while requiring only 0.14 additional questions per dialogue. Fixed-budget analyses show that the distinction persists when question count is controlled, while a matched expected-0/1-risk student control further isolates the contribution of asymmetric severity weighting. These results support moving proactive medical dialogue beyond uncertainty reduction toward consequence-aware evidence acquisition.
Figures & tables
Figure 1: Information value and diagnostic consequence can prioritize different questions. ESR evaluates candidate questions before their answers are observed and prioritizes evidence according to expected reduction in severity-aware terminal risk. The current patient’s hidden answer is not used during question ranking.
Method
Acc. ↑
Q
HSM ↓
Sev. ↓
Pop.-HS ↓
Cost ↓
Cap ↓
E-Entropy
.9123±.0015
2.182±.020
.0645±.0000
.0447±.0003
.0430±.0000
.2377±.0014
.0073±.0015
ESR
.9320±.0017
2.322±.025
.0455±.0048
.0333±.0015
.0303±.0032
.1886±.0008
.0177±.0038
Δ ( ESR − E-Entropy )
+.0197±.0025
+.140±.015
−.0190±.0048
−.0113±.0019
−.0127±.0032
−.0491±.0021
+.0103±.0051
Table 1: Main shared-stopping comparison across three seeds. Values are mean ± SD over seeds 42/43/44; only next-question supervision differs. Δ is the seed-wise ESR − E-Entropy difference.
Method
Acc. ↑
Q
HSM ↓
Sev. ↓
E-Entropy
.911
2.161
.0645
.0447
Exp.-0/1-Risk
.913
2.509
.0690
.0447
ESR
.931
2.294
.0435
.0337
Table 2: Frozen seed-42 student objective decomposition. All methods use expected-answer marginalization.
Large language models (LLMs) have made substantial progress on medical question-answering, yet effective medical dialogue also requires learning to ask questions that uncover relevant patient information. To train such dialogue policies, a common pipeline combines supervised fine-tuning with reinforcement learning (RL) based on final diagnostic correctness. However, this outcome-based supervision does not directly distinguish the contributions of individual questions and provides no question-level feedback for unexecuted alternatives. To address this gap, we introduce PCQC (Privileged Counterfactual Question Credit), which uses privileged patient information during training to learn from questions never asked. During training, PCQC makes alternative questions directly comparable at the same dialogue state by using privileged patient facts to construct their answers. A frozen diagnostic scorer evaluates the diagnostic utility of each resulting question-answer pair by how strongly it supports the correct diagnosis. PCQC turns these comparisons into relative question credit that teaches the policy which questions to favor, directly supervising both executed and unexecuted questions alongside outcome-based RL without requiring complete rollouts for the unexecuted alternatives. Extensive experiments across four medical benchmarks demonstrate that PCQC achieves 63.10% mean diagnostic accuracy, outperforming GRPO and ATPO by 4.38 and 4.21 percentage points, respectively. These gains are achieved with 33.1% fewer inquiry turns than GRPO.
Chenxuan Li, Jiayi Wan, Xinrong Chen +3
Peking University, Beijing, China · Department of Medical Bioinformatics, School of Basic Medical Sciences, Peking University, Beijing, China
Patient-facing medical LLMs and agents increasingly answer symptom questions before clinician contact, where the key safety question is what action the user should take next. We introduce CARE-Bench, a source-grounded benchmark that evaluates sequential patient-facing triage as a four-label per-turn current-action task. CARE-Bench contains 500 cases and 1,059 evaluated patient-disclosure prefixes reconstructed from medical dialogue, consultation, and follow-up-question sources. We evaluate 11 models on 269 held-out rounds under unprompted and minimally prompted open-ended protocols, using a fixed GPT-5.5 mapper to code each response into the four-label action space. Unprompted macro-F1 remains low, ranging from 31.2 to 50.4. Prompting improves 10 of 11 models, with prompted macro-F1 ranging from 46.9 to 63.4, but substantial threshold errors remain. Prompted models often recommend care before needed clarification is obtained; when the correct action was to ask for more information, only 33.5% of prompted outputs preserved the step. The persistence of these errors after prompting suggests that patient-facing triage is not a simple prompting problem and supports explicit evaluation of action timing before deployment.
Yining Hua, Hongbin Na, Cyrus Ayubcha
1Harvard University · University of Technology Sydney
Interactive medical questioning is essential in clinical consultations, where physicians must actively gather necessary patient information. Yet existing medical Large Language Models (LLMs) predominantly follow a reactive paradigm, risking diagnostic errors by answering before seeking sufficient details. To bridge this gap, we propose ProMed, a reinforcement learning framework that transitions LLMs toward a proactive paradigm, enabling them to ask clinically valuable questions before decision-making. Central to ProMed is the Shapley Information Gain (SIG) reward, which quantifies a question's clinical utility as the amount of newly acquired information, while considering its contextual importance via Shapley values. We integrate SIG into a two-stage training pipeline: (1) SIG-Guided Model Initialization uses Monte Carlo Tree Search to construct high-reward interaction trajectories for supervision, and (2) SIG-Augmented Policy Optimization, with a novel SIG-guided Reward Distribution Mechanism that prioritizes informative questions for fine-grained optimization. Experiments on partial-information medical benchmarks show that ProMed significantly outperforms state-of-the-art methods by 6.29% on average and delivers a 54.45% gain over the reactive paradigm, and generalizes robustly to out-of-domain cases. Our codes are available at https://github.com/hxxding/ProMed.
Hongxin Ding, Baixiang Huang, Yue Fang +8
National Engineering Research Center of Software Engineering, Peking University, China · School of Computing and Data Science, The University of Hong Kong · School of Computer Science, Peking University, Beijing, China +2