Organizations: Peking University, Beijing, China · Northeastern University, Boston, MA, USA · Sun Yat-sen University, Guangzhou, China · Department of Medical Bioinformatics, School of Basic Medical Sciences, Peking University, Beijing, China
Proactive medical dialogue requires an agent to decide what to ask from incomplete patient information. Existing information-seeking approaches commonly prioritize questions that most reduce diagnostic uncertainty, but this criterion overlooks an important property of medical diagnosis: different diagnostic errors can carry substantially different consequences. The most informative question may therefore differ from the one most valuable for the downstream decision. We propose Expected-Severity-Risk (ESR), a consequence-aware question-supervision objective that values each candidate by its expected reduction in severity-aware terminal risk. Because questions must be selected before their answers are observed, ESR marginalizes over possible answers using train-only population statistics. Its rankings are then distilled into a prefix-only language policy, requiring no teacher-side risk computation at deployment. Across three matched Qwen3-4B training seeds on DDxPlus, ESR reduces mean high-severity diagnostic miss from 0.0645 to 0.0455 (29.5% relative reduction) and improves mean diagnostic accuracy from 0.9123 to 0.9320 while requiring only 0.14 additional questions per dialogue. Fixed-budget analyses show that the distinction persists when question count is controlled, while a matched expected-0/1-risk student control further isolates the contribution of asymmetric severity weighting. These results support moving proactive medical dialogue beyond uncertainty reduction toward consequence-aware evidence acquisition.
Figures & tables
Figure 1: Information value and diagnostic consequence can prioritize different questions. ESR evaluates candidate questions before their answers are observed and prioritizes evidence according to expected reduction in severity-aware terminal risk. The current patient’s hidden answer is not used during question ranking.
Method
Acc. ↑
Q
HSM ↓
Sev. ↓
Pop.-HS ↓
Cost ↓
Cap ↓
E-Entropy
.9123±.0015
2.182±.020
.0645±.0000
.0447±.0003
.0430±.0000
.2377±.0014
.0073±.0015
ESR
.9320±.0017
2.322±.025
.0455±.0048
.0333±.0015
.0303±.0032
.1886±.0008
.0177±.0038
Δ ( ESR − E-Entropy )
+.0197±.0025
+.140±.015
−.0190±.0048
−.0113±.0019
−.0127±.0032
−.0491±.0021
+.0103±.0051
Table 1: Main shared-stopping comparison across three seeds. Values are mean ± SD over seeds 42/43/44; only next-question supervision differs. Δ is the seed-wise ESR − E-Entropy difference.
Method
Acc. ↑
Q
HSM ↓
Sev. ↓
E-Entropy
.911
2.161
.0645
.0447
Exp.-0/1-Risk
.913
2.509
.0690
.0447
ESR
.931
2.294
.0435
.0337
Table 2: Frozen seed-42 student objective decomposition. All methods use expected-answer marginalization.
National Engineering Research Center of Software Engineering, Peking University, China · School of Computing and Data Science, The University of Hong Kong · School of Computer Science, Peking University, Beijing, China +2