Reinforcement learning post-training for language models relies on two reward designs: human preferences (RLHF, DPO) and binary verifiers (RLVR). Clinical question answering fits neither. Near-correct answers differ by a single substituted entity, and no executable check decides clinical correctness. We instantiate a soft verifier from a maintained controlled vocabulary: UMLS Concept Unique Identifier overlap (via scispaCy, set-level F1) gives a graded, externally specified reward computed without a model in the loop. We combine it inside GRPO with an entropy-normalised LLM judge, which covers the safety and evidence axes overlap cannot see, and a small consistency penalty on padding and repetition that keeps early-training samples scorable. This three-term composite improves over SFT on Phi-3-mini (3.8B) over MedQA by 2.9% on EM (0.700 vs 0.680) and 39% on Token-F1 (0.202 vs 0.145); on Llama-3.2-3B the corresponding gains are 14% on EM and 35% on Token-F1. We report Token-F1 as the primary metric because it credits partially-correct clinical content that EM discards at this open-generation scale. Main-table results are means over 3 seeds with standard deviations below 0.005. The method transfers to PubMedQA, where training on the PubMedQA train set with the same composite reward improves Token-F1 over SFT by 22% on Phi-3-mini and 17% on Llama-3.2-3B without retuning. A reward ablation on Phi-3, varying the judge-ontology split at a fixed consistency weight, attributes 3 EM points to the ontology term, the contribution that catches entity substitutions the judge cannot. Three negative findings constrain the design: DPO under random negatives underperforms SFT for strong-prior models but helps the weakest-prior one; PPO under a sparse neural reward diverges; GRPO with KL-in-loss collapses at 7B.
Figures & tables
Figure 1. Composite-reward GRPO. The policy samples G=4 responses per question; each is scored by the frozen LLM judge, by UMLS CUI overlap against the reference, and by a consistency penalty on padding and repetition. The three terms are combined as in Eq. 1 ( α=0.6 , β=0.3 , γ=0.1 ). Group-normalised advantages update the LoRA adapter. Reference and judge remain frozen throughout training.
Model
Method
EM
Token F1
SeqM
Phi-3-mini (3.8B)
Zero-shot
0.560
0.070
0.132
SFT
0.680 ±0.004
0.145 ±0.002
0.183 ±0.003
DPO
0.600
0.143
0.183
GRPO (ours)
0.700 ±0.005
0.202 ±0.004
0.238 ±0.005
Qwen2.5-7B (7B)
Zero-shot
0.565
0.126
0.170
SFT
0.775
0.134
0.162
Table 1. MedQA test results (500 examples, greedy decoding). GRPO/SFT numbers are means over 3 seeds; standard deviations in subscript. DPO and Qwen2.5-7B SFT rows (no subscript) are single-seed runs. Bold = best per model per metric. Qwen2.5-7B GRPO collapsed across all configurations (§ 3.5 ).
Reward
EM
Token F1
SeqM
Judge only ( α=0.9 , β=0 )
0.667 ±0.004
0.162 ±0.003
0.201 ±0.004
UMLS only ( α=0 , β=0.9 )
0.592 ±0.007
0.143 ±0.005
0.185 ±0.006
Composite ( α=0.6 , β=0.3 )
0.700 ±0.005
0.202 ±0.004
0.238 ±0.005
Table 2. Reward component ablation on Phi-3-mini, means over 3 seeds
Test-time reinforcement learning adapts a model on its own unlabeled test set using majority-vote pseudo-labels and has shown strong results in mathematics. We show that this recipe collapses on medical multiple-choice QA: accuracy stagnates while output diversity rapidly declines. Through a controlled experiment that keeps the questions, model, and optimizer fixed while changing only the answer space, we trace this failure to answer-space structure rather than domain difficulty. In small answer spaces, incorrect rollouts often collide on the same wrong pseudo-label and reinforce it; in large answer spaces, they disperse and receive little reward. This diagnosis motivates PROSE, Process Reward Guided Self-Training, which rewards reasoning quality instead of answer agreement. PROSE scores each reasoning step with a medical process reward model, assigns the trajectory reward as the minimum score across steps, and enforces answer-format constraints. Without labels, PROSE substantially improves a general Llama model, surpassing purpose-built medical models and matching much larger systems. Because the process signal is internalized into the policy, the adapted model requires no reward model at inference and transfers its gains to unseen datasets. We further show that the minimum aggregation is essential: mean aggregation can be exploited, saturating the proxy reward while degrading accuracy.
Kailong Fan, Anqi Pu, Yichen Wu +7
1Harvard Medical School/MGH · 3Harvard University · 2Zhejiang University
Reinforcement Fine-Tuning (RFT) has enabled medical Multimodal Large Language Models (MLLMs) to produce Chain-of-Thought (CoT) reasoning for visual question answering, yet these models suffer from confidence miscalibration---a systematic gap between expressed certainty and actual diagnostic accuracy that undermines clinical trust. We propose CARE, a Confidence-Aware medical REasoning framework that jointly optimizes accuracy and calibration through a dual-stage pipeline. First, a scalable Medical-CoT synthesis provides structured cold-start data for Supervised Fine-Tuning. Second, Group Relative Policy Optimization (GRPO) with a novel Confidence-Aware Reward (CAR) mechanism ties the model's confidence to diagnostic correctness within the reward signal. Across three Medical VQA benchmarks, CARE achieves the highest diagnostic accuracy while obtaining the lowest Expected Calibration Error and Hallucination Rate, establishing a foundation for trustworthy clinical decision support. Our code is available at https://github.com/anotherbricki/CARE.
Yuetian Du, Yucheng Wang, Zhenyuan Chen +9
Zhejiang University · Ant Group · University of Michigan +1
Large language models (LLMs) have made substantial progress on medical question-answering, yet effective medical dialogue also requires learning to ask questions that uncover relevant patient information. To train such dialogue policies, a common pipeline combines supervised fine-tuning with reinforcement learning (RL) based on final diagnostic correctness. However, this outcome-based supervision does not directly distinguish the contributions of individual questions and provides no question-level feedback for unexecuted alternatives. To address this gap, we introduce PCQC (Privileged Counterfactual Question Credit), which uses privileged patient information during training to learn from questions never asked. During training, PCQC makes alternative questions directly comparable at the same dialogue state by using privileged patient facts to construct their answers. A frozen diagnostic scorer evaluates the diagnostic utility of each resulting question-answer pair by how strongly it supports the correct diagnosis. PCQC turns these comparisons into relative question credit that teaches the policy which questions to favor, directly supervising both executed and unexecuted questions alongside outcome-based RL without requiring complete rollouts for the unexecuted alternatives. Extensive experiments across four medical benchmarks demonstrate that PCQC achieves 63.10% mean diagnostic accuracy, outperforming GRPO and ATPO by 4.38 and 4.21 percentage points, respectively. These gains are achieved with 33.1% fewer inquiry turns than GRPO.
Chenxuan Li, Jiayi Wan, Xinrong Chen +3
Peking University, Beijing, China · Department of Medical Bioinformatics, School of Basic Medical Sciences, Peking University, Beijing, China