Unmentioned Checklist Findings Change How Reinforcement Learning Appears to Improve Chest Radiograph Report Checking
Authors: Ali Vosoughi, Akhil Kasturi, Chenliang Xu, Axel Wismueller
Organizations: Department of Computer Science, University of Rochester, Rochester, NY 14627, USA. · Department of Electrical and Computer Engineering, University of Rochester, Rochester, NY 14627, USA. · Department of Imaging Sciences, University of Rochester Medical Center, Rochester, NY 14642, USA. · Department of Biomedical Engineering, University of Rochester, Rochester, NY 14627, USA.
Automated checks of radiology reports may rely on AI-generated checklists that leave findings unmentioned. We used reinforcement learning to train a vision-language model to fill in a 12-finding checklist from a chest radiograph without seeing the sentence under test; a separate checking model judged the sentence from the checklist. On held-out patients, a rule-based check and an independent medical checker, neither used in training, measured discrimination gains (Youden index) of 12.6% and 11.8%; only the rule-based check met the prespecified false-alarm criterion. Switching to the training format, which fixes finding order and enters unmentioned findings as absent, raised the training checker's measured gain and lowered the independent checker's, a prespecified comparison that yielded 6.2% (95% interval 2.0% to 10.5%) and, post hoc on held-out patients, 7.7%. Across 8 checking models, acceptance of a label-consistent negative statement about an unmentioned finding ranged from 1.0% to 97.0%. Labels were report-derived, not radiologist-adjudicated.
Figures & tables
Fig. 1: Does “not mentioned” mean “absent”? Without seeing the report sentence, an AI model fills in a checklist of 12 possible findings from 1 radiograph, shown here in short with its free-text note for 1 illustrative CheXpert Plus study. The checklist has no entry for pleural effusion, although its general note says “No other abnormalities are visible in this radiograph.” The sentence “No pleural effusions.” agrees with the study’s report-derived label. Both checking models reject the sentence on the checklist as generated and accept it in the training format, which puts the findings in a fixed order and enters unmentioned findings as absent. Because the general note can itself be read as excluding effusion, the example shows a missing finding-specific entry rather than complete silence. Across 237 matched validation pairs, switching to the training format moves the gain from training in pair-level discrimination ( Y ) from 3.8% to 6.0% for the checker used in training and from 7.0% to 3.1% for an independent medical checker.
Fig. 2: Study design. (a) Training. A matched pair is 2 frontal studies whose report-derived labels agree on 11 findings and differ on the finding a report sentence describes; here the sentence “No pleural effusion.” is consistent with the label of the source study (effusion absent) and inconsistent with the label of the matched study (effusion present). Each image is processed separately. The checklist model sees only the image and fills in a 12-finding checklist, which is rewritten in the training format, a fixed finding order with every unmentioned finding entered as absent. A frozen checking model reads that checklist and the sentence, never the image, and the checklist model is rewarded when the verdict agrees with the report-derived label (reinforcement learning by group relative policy optimization, GRPO). Pairs are matched on report-derived labels, not on radiographic appearance. Burned-in technologist initials on the 2 radiographs were masked; side and position markers were kept. The rows below the images show the effusion entry of this pair’s checklists before training and after training in run 1; the pair was chosen by a rule fixed in advance, the first effusion pair whose trained checklists agree with the report-derived labels on both images under the rule-based check and the independent checker. A selected example illustrates the mechanism and does not show typical improvement. (b) Evaluation. The saved checklists of the untrained and trained models are read as generated and in the training format, with 2 intermediate formats that separate the fixed order from the entered absences, by the checking model used in training, by an independent medical checking model never used in training, and by a rule-based check of the present or absent mark that uses no language model.
Term
Meaning in this paper
Checklist model
The vision-language model that looks at 1 frontal radiograph and fills in the checklist. It never sees the report sentence under test and is the only model that is trained.
Checklist
Up to 12 listed findings, each entry giving present or absent, side, an optional box, and a short note. Findings may be left unmentioned.
Checker
Any of the 3 methods that give a verdict on a sentence from a checklist: the training checker, the independent checker, and the rule-based check.
Checking model
A frozen language model that reads only the checklist and 1 report sentence and answers supported or unsupported. The training checker supplied the reward; the independent checker, MedGemma, never did.
Rule-based check
Supports the sentence when the checklist’s present or absent mark for the sentence’s finding agrees with the sentence. It reads the repaired checklist, counting an unmentioned finding as absent, and uses no language model.
Report-derived label
Present or absent for each finding, extracted automatically from the report text by the CheXpert labeler, not a radiologist’s reading of the image.
Table 1: Terms used in this paper. Every measure is computed against report-derived labels, not radiologist reads of the images.
Checker
Untrained Y
Untrained FA
ΔY [95%]
Δ FA (upper 95%)
Gain
False alarms
Validation pairs, 237
Rule-based check
3.0
45.1
+ 8.3 [ + 2.0, + 14.1]
+ 1.7 ( + 6.6)
met
not met
Independent checker
2.1
36.3
+ 7.0 [ + 0.4, + 13.3]
+ 4.4 ( + 9.6)
met
not met
Training checker
4.6
46.0
+ 3.8 [ − 1.0, + 8.2]
+ 12.0 ( + 15.9)
not in criterion
Held-out test pairs, 273
Rule-based check
3.3
40.7
+ 12.6 [ + 6.9, + 18.2]
− 2.7 ( + 1.9)
met
met
Table 2: Training improved discrimination for both checkers not used in training, but the joint criterion on discrimination and false alarms failed in both populations. Replication on the validation pairs and confirmation on held-out test pairs from unseen patients, all entries in percent; checklists read as generated. The protocol was recorded after the 3 training jobs were launched and before any of their checklists was read; the runs differ only in their random seed. ΔY is the run-averaged paired change in discrimination from the untrained model, with its two-sided 95% interval over patient groups (per-run values in the Supplementary Information), and Δ FA the averaged change in the false-alarm rate with its one-sided 95% upper bound. The criterion asked each checker not used in training for ΔY≥5% with the lower bound above zero (gain) and an upper bound on Δ FA of at most 5% (false alarms). A change in a rate is the difference of the 2 percentages. Intervals are conditional on the 3 trained models.
Training checker
Independent checker
Format of untrained and trained checklists
ΔY
Δ FA
ΔY
Δ FA
As generated (evaluation default)
+ 3.8 [ − 1.0, + 8.2]
+ 12.0 [ + 7.6, + 16.5]
+ 7.0 [ + 0.4, + 13.3]
+ 4.4 [ − 1.4, + 10.5]
Fixed order, unmentioned left out
+ 3.7 [ − 1.0, + 8.2]
+ 12.0 [ + 7.4, + 16.6]
+ 6.2 [ − 0.3, + 12.8]
+ 4.9 [ − 0.9, + 10.9]
Model’s order, absences entered
+ 5.8 [ + 0.8, + 10.8]
+ 6.4 [ + 1.4, + 11.5]
+ 5.4 [ − 1.0, + 12.0]
+ 0.8 [ − 4.5, + 6.5]
Training format (fixed order, absences entered)
+ 6.0 [ + 1.0, + 10.8]
+ 6.9 [ + 1.9, + 11.9]
+ 3.1 [ − 3.7, + 9.5]
+ 1.8 [ − 3.9, + 7.9]
Checker difference, training format (prespecified)
+ 6.2 [ + 2.0, + 10.5] on 237 pairs
Table 3: The 2 checkers respond in opposite directions to the training format. The same saved validation checklists of the 3 replication runs, read in 4 formats, all entries in percent. Each row applies 1 format to untrained and trained checklists alike, with run-averaged changes in discrimination ( ΔY ) and the false-alarm rate ( Δ FA) and 95% intervals over patient groups, on 237 pairs for the as-generated and training-format rows and 236 for the 2 intermediate rows. The difference between checkers is the independent checker’s change from the training format to the checklists as generated minus the training checker’s; it was prespecified before any replicated checklist was read. Its split into the fixed order and the entered absences (the 2 middle rows and the last row) was declared after the first run’s checklists had been read as generated and in the training format, before the 2 intermediate formats were read, and is post hoc. The rule-based check is unchanged by every row.
Negative accepted
Difference from the training checker
Checker
(finding unmentioned)
Validation
Held-out
Training checker
64.0 / 53.5
–
–
MedGemma-4B
79.0 / 74.0
+ 6.2 [ + 2.0, + 10.5]
+ 7.7 [ + 2.8, + 12.8]
Gemma-3-4B
1.0 / 1.6
+ 0.6 [ − 4.0, + 4.8]
+ 4.9 [ − 0.2, + 10.2]
Gemma-3-12B
91.0 / 91.3
+ 1.5 [ − 1.6, + 4.7]
+ 3.2 [ + 0.1, + 6.3]
Gemma-3-27B
56.0 / 50.4
+ 0.3 [ − 2.4, + 3.0]
+ 1.9 [ − 1.1, + 5.0]
Table 4: Checking models disagree on whether an unmentioned finding counts as absent. All entries in percent. Negative accepted: verdicts accepting a negative statement that agrees with its report-derived label, such as no pleural effusion, about a finding the checklist does not mention (100 validation / 127 held-out test verdicts). Difference: the checker’s format effect minus the training checker’s (Table 3 ), 95% intervals over patient groups, on 221 to 274 pairs; in the last 2 rows both primary checkers read under an instruction to quote the checklist entry or to reason first. MedGemma’s differences are those of Table 3 ; all other entries come from post hoc rereads of the same saved checklists. Phi-4-mini, OLMo-2-7B, and Qwen3-VL-30B-A3B, which complete the 8 checkers, are reported in the Supplementary Information.
Fig. 3: What each checker measures, and whether the trained checklists describe the image. 3 replication runs, in percent. (a) The run-averaged improvement in discrimination ( Y ) from untrained to trained checklists when the same saved checklists are read as generated and in the training format, on 237 validation pairs (prespecified) and 270 held-out test pairs (post hoc). Below, the independent checker’s change from the training format to the checklists as generated minus the training checker’s, with its 95% interval over patient groups; neither checker’s own change is resolved on the validation pairs. The held-out estimates differ slightly from Table 2 because these rereads use the 270 pairs on which both checkers returned a verdict in every format rather than the 273 endpoint pairs. (b) Descriptive: the rule-based check’s Y on the 238 validation pairs for the untrained model and each run (squares) beside the full range of 200 shuffles of the checklists among images of the same finding and sentence direction (grey bars, dark tick at the maximum; not confidence intervals). A value right of the bar exceeds every shuffle.
Appendix figures & tables17 assets
Supplementary material from the paper’s appendix.
Appendix
Training judge
Independent judge
Rule-based reader
Record
Y
FS
Y
FS
Y
FS
Untrained writer
4.6
45.8
2.1
36.3
2.9
45.0
Development run, step 40
8.8
53.8
12.2
35.9
12.6
44.1
Change
+ 4.2
+ 8.0
+ 10.1
− 0.4
+ 9.7
− 0.8
Labels as flags
68.5
27.7
53.4
4.2
100.0
0.0
Appendix
Table S1: The development run’s records read in their original form by 3 readers on the 238 validation pairs, all entries in percent (237 common pairs for the independent judge’s paired untrained-to-trained comparison, since its output for 1 trained record contained no parseable verdict, while the label row uses all 238). Y is the pair-level Youden index and FS the false-strike rate. Changes are paired, with 95% intervals of [−1.7%,+10.6%] , [+2.1%,+18.1%] , and [+3.3%,+16.5%] for Y and [+4.1%,+12.1%] , [−5.6%,+4.8%] , and [−5.5%,+3.8%] for FS. Records written from the labels show that each reader has room to register a perfect flag record, and that the judges fall short of it even then, as they also assess sentence details that the flags do not specify.
Training judge
Independent judge
Records read
ΔY
Δ FS
ΔY
Δ FS
Original form
+ 4.2 [ − 1.7, + 10.6]
+ 8.0 [ + 4.1, + 12.1]
+ 10.1 [ + 2.1, + 18.1]
− 0.4 [ − 5.6, + 4.8]
Original form, sides not applicable
+ 3.8 [ − 2.4, + 10.2]
+ 7.6 [ + 3.7, + 11.5]
+ 11.4 [ + 3.4, + 19.2]
− 0.4 [ − 5.5, + 4.6]
Reward-scoring form, every field
+ 6.7 [ + 0.4, + 13.2]
+ 6.3 [ + 2.1, + 10.4]
+ 6.3 [ − 1.3, + 13.6]
− 0.4 [ − 5.0, + 4.3]
Reward-scoring form, notes removed
+ 7.6 [ + 1.7, + 13.6]
+ 3.4 [ − 0.8, + 7.6]
+ 5.9 [ − 1.7, + 13.7]
− 2.5 [ − 8.0, + 2.9]
Reward-scoring form, flags only
+ 8.0 [ + 1.7, + 14.5]
+ 2.9 [ − 1.7, + 7.5]
+ 3.4 [ − 3.0, + 10.1]
+ 3.4 [ − 1.7, + 8.2]
Appendix
Table S2: The development run’s saved records read in different forms on the 238 validation pairs, all entries in percent (237 for the independent judge), each form applied to both the untrained and the trained records. Entries are paired changes in the pair-level Youden index ( ΔY ) and the false-strike rate ( Δ FS) with 95% intervals over patient groups.
Development run
First registered run
Finding
images / groups / informative
images / groups / informative
Pleural effusion
252 / 265 / 118 (45)
256 / 273 / 124 (45)
Edema
229 / 229 / 110 (48)
283 / 283 / 120 (42)
Consolidation
200 / 211 / 68 (32)
200 / 213 / 61 (29)
Cardiomegaly
140 / 140 / 46 (33)
158 / 158 / 24 (15)
Atelectasis
132 / 132 / 21 (16)
167 / 167 / 30 (18)
Appendix
Table S3: Training exposure per finding, from the per-sentence reward logs of the training images, shares in parentheses in percent. A group is a single image’s sentence over its 16 records, and a group is informative when its mean reward lies strictly between 0 and 1, the only groups that carry a verdict-reward signal.
Item
Setting
Writer and judge
Qwen3-VL-8B-Instruct, bfloat16; judge frozen and served separately
Adapter
low-rank adaptation (LoRA) rank 16, α=32 , dropout 0, on q, k, v, o, gate, up, down projections of the language model; vision tower untouched
Algorithm
group relative policy optimization, Kullback-Leibler (KL) loss to the initial writer (coefficient 0.01, low-variance estimator), no entropy bonus
Dynamic sampling
groups with no reward variation discarded and the batch refilled
Rollouts
16 records per image, temperature 1.0, up to 1,024 new tokens
Batch
8 images per step, mini-batch 4, learning rate 10−4 , 40 steps, adapters saved every 10 steps
Appendix
Table S4: Configuration of the registered replication, identical to the learning-rate run except for the seeds.
Figure S1: Training logs of the 3 replication runs with the standard reward (black) and the 3 runs with the omission penalty (ochre), runs as line styles, each curve a centred 5-step running mean of the logged steps. (a) Mean record length in tokens. (b) The Kullback-Leibler divergence term to the initial writer in the training loss.
Figure S2: Registered endpoints for each training run and their mean, records read in their original form, in percent. (a) Change in the pair-level Youden index Y from the untrained writer, with 95% intervals over patient groups. (b) Change in the false-strike rate, each run with its 95% interval and the mean with its one-sided 95% upper bound. The red dashed line marks the registered 5% gain criterion in (a) and the 5% false-strike margin in (b), and the 237 validation and 273 held-out test pairs are those common to all endpoint reads.
Validation, 237 pairs
Held-out test, 273 pairs
Reader
Writer
ΔY
Δ FS
ΔY
Δ FS
Rule-based reader
untrained Y , FS
3.0
45.1
3.3
40.7
run 1
+ 11.0 [ + 3.4, + 18.6]
+ 0.4
+ 13.6 [ + 7.0, + 20.1]
− 4.0
run 2
+ 9.3 [ + 2.5, + 16.0]
+ 3.0
+ 11.0 [ + 4.5, + 17.4]
− 2.6
run 3
+ 4.6 [ − 1.7, + 11.0]
+ 1.7
+ 13.2 [ + 6.1, + 20.7]
− 1.5
mean
+ 8.3 [ + 2.0, + 14.1]
+ 1.7
+ 12.6 [ + 6.9, + 18.2]
− 2.7
Appendix
Table S5: Registered endpoints of the replication for each training seed, all entries in percent, on the pairs common to all its endpoint reads, with records read in their original form. Each reader’s first row gives the untrained writer’s pair-level Youden index Y and false-strike rate FS, and the other rows the paired change from untrained to trained records, with 95% intervals over patient groups for ΔY . The mean rows are the seed-averaged changes of the main paper’s Table 2.
Figure S3: What the writers record, on the validation records, descriptive. (a) For each finding, the share of images whose report label is positive (detection) or negative (false-positive flags) on which the target flag is set, for the untrained writer (open squares) and the mean of the 3 replication runs (filled squares), with the pairs per finding in parentheses, in percent. (b) How many of the 12 findings each writer lists in a record, as the share of its 476 records in each count, with the mean and the share of records that omit at least 1 finding at the right. Listing all 12 findings measures completeness, not correctness.
Sentence (pairs)
Judge
Original form
Not assessed
Absent
Marking
Assertion
Denies the finding (148)
Training
+ 15.8
+ 20.0
+ 10.6
− 4.3 [ − 6.8, − 2.0]
+ 9.5 [ + 6.0, + 12.7]
Independent
+ 13.3
+ 9.7
+ 8.3
+ 3.6 [ + 0.5, + 6.6]
+ 1.4 [ − 1.3, + 4.0]
Asserts the finding (88)
Training
+ 5.7
+ 0.8
− 0.8
+ 4.9 [ + 0.8, + 9.5]
+ 1.5 [ − 1.1, + 4.9]
Independent
− 11.0
− 11.4
− 11.7
+ 0.4 [ − 5.7, + 5.4]
+ 0.4 [ − 4.1, + 4.5]
Appendix
Table S6: Exploratory split of the not-assessed contrast by sentence direction, all entries in percent, on the 236 validation pairs common to the original, not-assessed, and absent forms in the writer’s order. Entries in the first 3 columns are the seed-averaged change in the false-strike rate from untrained to trained records with the omitted findings left out (in the original form), marked not assessed, or stated absent, all in the writer’s order. Marking is original minus not assessed, and assertion is not assessed minus absent, with 95% intervals over patient groups, where a positive value is a reduction of the false-strike increase.
Training judge
Independent judge
Form of untrained and trained held-out test records
ΔY
Δ FS
ΔY
Δ FS
Original form
+ 6.4 [ + 2.1, + 10.8]
+ 7.2 [ + 2.2, + 11.9]
+ 12.5 [ + 6.0, + 18.7]
+ 3.8 [ − 2.3, + 9.5]
Fixed order, omissions left out
+ 6.7
+ 6.7
+ 11.0
+ 4.7
Writer’s order, omissions absent
+ 11.2
+ 1.9
+ 10.5
− 0.5
Reward-scoring form
+ 11.2 [ + 6.8, + 16.1]
+ 2.0 [ − 3.2, + 6.9]
+ 9.6 [ + 3.7, + 15.6]
− 1.0 [ − 7.0, + 4.4]
Appendix
Table S7: Post hoc re-evaluation of the saved held-out test records of the untrained writer and the 3 replication runs, all entries in percent, in the forms of the validation decomposition, with seed-averaged paired changes in the pair-level Youden index ( ΔY ) and the false-strike rate ( Δ FS), declared before any of these forms was read, on the 270 held-out test pairs where every form parses. The difference between the judges’ responses to the whole change of form is 7.7%[2.8%,12.8%] , its fill component 6.4%[2.2%,10.6%] , and its order component an unresolved 1.3%[−1.1%,3.8%] . Stating the omissions as absent lowers both judges’ false-strike increase by 5.0%[2.6%,7.3%] . The registered held-out test endpoints, on records in the original form, are those of the main paper’s Table 2.
Judge
Pairs
ΔY original
ΔY reward-scoring
Form effect, ΔY
Fill effect, Δ FS
Minus training judge
Validation pairs
Qwen3-VL-8B (training)
236
+ 3.8
+ 6.1
− 2.3 [ − 5.2, + 0.6]
+ 5.4 [ + 2.6, + 8.0]
MedGemma-4B (independent)
236
+ 7.2
+ 3.2
+ 4.0 [ − 0.1, + 8.0]
+ 3.3 [ + 0.6, + 6.2]
+ 6.2 [ + 2.0, + 10.6]
Qwen3-VL-30B-A3B
238
+ 3.9
+ 6.3
− 2.4 [ − 5.6, + 1.2]
+ 6.3 [ + 3.3, + 9.0]
− 0.1 [ − 2.7, + 2.4]
Phi-4-mini
238
+ 0.6
+ 1.0
− 0.4 [ − 4.8, + 4.0]
+ 2.2 [ − 0.3, + 4.5]
+ 1.8 [ − 2.4, + 6.1]
OLMo-2-7B
238
+ 4.2
+ 5.7
− 1.5 [ − 5.8, + 2.8]
+ 5.7 [ + 2.5, + 8.6]
+ 0.7 [ − 3.5, + 4.7]
Appendix
Table S8: The record-form effect under 5 judges reading the same saved records of the untrained writer and the 3 replication runs, all metric values in percent. ΔY is the run-averaged paired change in the pair-level Youden index from untrained to trained records, and the form effect is its value in the original form minus its value in the reward-scoring form, with 95% intervals over patient groups, so a negative form effect on ΔY means the judge registers more of the learning in the reward-scoring form. The fill effect on Δ FS is the seed-averaged change in the false-strike rate with omitted findings left out minus with them written absent, averaged over the 2 orders, so a positive value means that writing the absences lowers the rise in false strikes. The last column is each judge’s form effect on ΔY minus the training judge’s, recomputed on that row’s pairs and the same draws, so it need not equal the difference from the training row printed above it. The judges are Qwen3-VL-8B-Instruct, MedGemma-4B-it, Qwen3-VL-30B-A3B-Instruct, Phi-4-mini-instruct, and OLMo-2-1124-7B-Instruct. Pairs are those on which every read of that comparison parses. The 3 further judges were added post hoc, declared before any of their reads, and all rows use 1 analysis script.
Figure S4: The 5 judges reading the same saved records with their default instruction (open markers) and with 1 added sentence stating that any unlisted finding is absent (filled markers), in percent. (a) Form effect on the gain, the change in Y in the original form minus that in the reward-scoring form. (b) How much writing the omitted findings as absent lowers the rise in false strikes, averaged over the 2 orders. Intervals and pair counts are in Table S9 .
Form effect on ΔY
Fill effect on Δ FS
Judge
Pairs
Default
Convention
Change
Default
Convention
Change
Validation pairs
Qwen3-VL-8B (training)
238
− 2.2
− 1.7
+ 0.6 [ − 1.6, + 2.5]
+ 5.3
+ 1.6
− 3.7 [ − 5.3, − 2.0]
Qwen3-VL-30B-A3B
238
− 2.4
− 2.8
− 0.4 [ − 2.6, + 1.7]
+ 6.3
+ 6.5
+ 0.2 [ − 1.2, + 2.0]
MedGemma-4B (independent)
234
+ 3.6
+ 5.7
+ 2.1 [ − 0.8, + 5.5]
+ 3.3
+ 2.5
− 0.8 [ − 2.3, + 0.6]
OLMo-2-7B
238
− 1.5
− 3.6
− 2.1 [ − 5.4, + 1.6]
+ 5.7
+ 6.4
+ 0.7 [ − 1.0, + 2.3]
Appendix
Table S9: The same saved records read by each judge with its default instruction and with 1 added sentence stating that any of the 12 findings the record does not list is absent, all metric values in percent, a post hoc control declared before any of its reads. Form and fill effects are defined as in Table S8 , recomputed here on the pairs where every read of the judge under both instructions parses (the between-judge rows use the pairs where both judges parse under the convention), and the change is the value under the convention minus the value under the default instruction, with 95% intervals over patient groups.
Comparison
Population
Pairs
Form effect, judge B
Form effect, judge A
B minus A
Gemma-3-4B vs training judge
validation
238
− 1.7 [ − 5.8, + 2.5]
− 2.2 [ − 5.0, + 0.7]
+ 0.6 [ − 4.0, + 4.8]
Gemma-3-4B vs training judge
held-out test
274
+ 0.1 [ − 4.9, + 5.1]
− 4.7 [ − 8.1, − 1.7]
+ 4.9 [ − 0.2, + 10.2]
Gemma-3-12B vs training judge
validation
238
− 0.7 [ − 4.1, + 2.8]
− 2.2 [ − 5.0, + 0.7]
+ 1.5 [ − 1.6, + 4.7]
Gemma-3-12B vs training judge
held-out test
274
− 1.6 [ − 4.0, + 0.6]
− 4.7 [ − 8.1, − 1.7]
+ 3.2 [ + 0.1, + 6.3]
Gemma-3-27B vs training judge
validation
238
− 2.0 [ − 5.2, + 1.3]
− 2.2 [ − 5.0, + 0.7]
+ 0.3 [ − 2.4, + 3.0]
Gemma-3-27B vs training judge
held-out test
274
− 2.8 [ − 6.1, + 0.5]
− 4.7 [ − 8.1, − 1.7]
+ 1.9 [ − 1.1, + 5.0]
Appendix
Table S10: Judge size and judge instruction, in percent, all reads declared post hoc before they were made. The form effect is ΔY in the original form minus ΔY in the reward-scoring form, each ΔY comparing trained with untrained records, as in Table S8 , with 95% intervals over patient groups. In the instruction rows judge A reads with its default instruction and judge B with the named one. Pair counts differ because a pair enters only when every read of both judges returns a supported or unsupported verdict on both of its images.
Target finding not mentioned by the record
Target finding listed
Judge
Population
Correct negative
Incorrect negative
Correct positive
Incorrect positive
Correct negative
Incorrect negative
Training judge
validation
64.0 → 99.0
65.8 → 97.5
6.6 → 6.6
4.8 → 8.3
64.1 → 64.1
52.4 → 52.6
MedGemma-4B
validation
79.0 → 99.0
71.8 → 97.4
15.8 → 31.6
24.1 → 39.8
67.9 → 72.6
56.3 → 61.3
Qwen3-VL-30B
validation
40.0 → 99.0
44.3 → 96.2
6.6 → 6.6
4.8 → 8.3
63.5 → 63.7
52.4 → 52.6
Phi-4-mini
validation
97.0 → 99.0
83.5 → 97.5
17.1 → 14.5
17.9 → 16.7
86.7 → 91.3
81.6 → 88.2
OLMo-2-7B
validation
45.0 → 97.0
49.4 → 89.9
3.9 → 18.4
3.6 → 14.3
64.9 → 65.1
53.6 → 54.4
Appendix
Table S11: What each judge does with silence, default instruction, in percent: the share of verdicts accepting the report sentence on the record as written → with every omission written absent, for the same saved records of the untrained writer and the 3 replication runs. Negative and positive denote the target finding’s absence and presence as assigned by the pair’s report-derived labels, correct and incorrect denote agreement with those labels, and the target finding is the sentence’s. Each verdict is weighted equally, pooled over the 4 record sets, and a parsed verdict that is neither supported nor unsupported counts as not accepted. Verdicts per cell, in column order, for the training judge (other judges at most 1 fewer through unparsed replies): validation 100, 79, 76, 84, 496, 517; held-out test 127, 130, 54, 62, 625, 622. Differences with 95% intervals over patient groups are in the released result files.
Judge
Form
Verdict and reason
Training
in the original form
rejected, “The record does not mention pleural effusions, so the sentence cannot be confirmed as true based on available data.”
Training
reward-scoring form
accepted, “Pleural Effusion is marked as absent in the record.”
Independent
in the original form
rejected, “The record does not mention pleural effusions.”
Independent
reward-scoring form
accepted, “The record indicates that Pleural Effusion is not present.”
Appendix
Table S12: The 2 judges on the record above in its 2 forms. The rule-based reader, which normalizes every record, accepts the sentence in either form.
Y , older / newer writer
Omission form
Encoding
Comprehension
Judge
blank
absent
absent − blank
sparse − complete
complete / sparse
Qwen3-VL-8B (training judge’s model)
90.0 / 88.8
91.6 / 89.4
− 1.0 [ − 2.5, + 0.6]
− 0.1 [ − 1.3, + 1.1]
100 / 100
Gemma-3-12B
89.7 / 87.4
91.8 / 89.1
− 0.3 [ − 2.2, + 1.6]
+ 0.1 [ − 1.1, + 1.3]
100 / 100
Qwen3-VL-30B-A3B
88.3 / 88.0
91.9 / 89.7
− 1.9 [ − 3.3, − 0.4]
+ 0.3 [ − 0.7, + 1.2]
100 / 97.9
Phi-4-mini
85.8 / 85.5
63.8 / 69.3
+ 5.8 [ + 3.0, + 8.5]
− 6.8 [ − 9.7, − 3.9]
87.5 / 100
OLMo-2-7B
61.6 / 60.5
60.1 / 56.5
− 2.5 [ − 5.7, + 0.6]
− 2.8 [ − 6.0, + 0.3]
87.5 / 91.7
Appendix
Table S13: The photograph study, read once under the frozen protocol, all entries in percent; study outcomes use the 510 writer-complete pairs and comprehension 48 synthetic records per encoding. Y is the pair-level Youden index of each untrained writer’s sparse-style records under each judge with omitted objects left blank or written as absent. The omission-form column gives the change in the writer difference, newer minus older, from blank to absent, and the encoding column the change from the complete state list to the present objects alone under the stated convention, with 95% intervals over 10,000 resamples of pairs. Comprehension is the share of 48 synthetic records per encoding read as the convention implies, with 90% required; the 2 judges below it are secondary and reported, not replaced. The last row holds the 2 registered estimands, each requiring an interval that excludes zero and an absolute estimate of at least 5%, computed from unrounded values.
Vision-language models that answer questions about chest radiographs are evaluated by their accuracy on labels derived from radiology reports. High benchmark accuracy is often interpreted as evidence that the model uses the image. A model that answers from the finding named in the question can score as well as a model that uses the radiograph. Keeping the question fixed, we audit eight open-weight systems by swapping in another patient's radiograph with the same or the opposite label, occluding the radiologist-marked region or an equal region elsewhere, and removing the radiograph or replacing it with noise or a photograph. On 2,548 yes-or-no questions from MIMIC-CXR, one multimodal model answers Yes regardless of the image, another multimodal model changes its answers without following the label, and four systems use the image but keep about half of their correct answers when the radiograph is swapped for an opposite-label radiograph. A medical model that receives only the question text scores 55.3% on the pooled questions, higher than two multimodal systems. It scores 91.8% where every finding is present, and answering Yes to every question scores 100% there. Where the image is necessary, the best multimodal system exceeds this model by 10.4% in balanced accuracy. The categories are unchanged on CheXpert. Confidence is not higher when a correct answer depends on the marked region. In a reader study with three radiologists, the two radiologists who read a balanced set of 200 cases score 86.0% and 82.0%, and the systems score 50.0% to 73.0%. Accuracy does not establish image use, but an intervention on the image can test it.
Mahshad Lotfinia, Sebastian Ziegelmayer, Lisa Adams +4
Pattern Recognition Lab, Friedrich-Alexander-Universität Erlangen-Nürnberg, Erlangen, Germany · Department of Diagnostic and Interventional Radiology, TUM University Clinic, School of Medicine and Health, Klinikum rechts der Isar, Technical University of Munich, Munich, Germany · Lab for AI in Medicine, RWTH Aachen University, Aachen, Germany +1
Vision-Language Models (VLMs) for radiology report generation are typically trained on retrospective clinical reports, which suffer from omission noise: clinically present findings are left unreported due to the omission of subtle findings. For example, prior studies show that cardiomegaly may be omitted from ICU chest X-ray reports when the imaging request is focused on monitoring support device placement. As a result, models trained with standard approaches inherit these omissions, learning to under-report findings themselves. We propose PU-DPO, a preference optimization framework to prevent omission noise from corrupting the preference signal. We reformulate the objective under a positive-unlabeled (PU) learning framework, treating absent mentions as unlabeled rather than truly negative. Our framework provides preference supervision using constructed contrastive pairs, generated using edits to model responses, producing variants that explicitly mention or omit a specific finding. Generated responses that mention the finding are naturally preferred in the context of visual evidence. Across semi-synthetic experiments and analyses on real-world chest radiograph benchmarks where adjudicated labels are available, PU-DPO yields consistent gains in detection rates and recovery of hidden positives across multiple pathologies, and is more robust to omission noise than prior approaches.
Yuta Kobayashi, Pradyun Ramesh, Muhammad Ahmed Chaudhry +5
Columbia University · Stanford University · Emory University
Vision-language models return structured chest-radiograph findings through interfaces exposing no confidence score, so a receiving institution cannot read off how far to trust an individual judgment. Whether agreement with an institution's reference standard transfers across sites, findings, prediction directions and question formats is largely unmeasured. We evaluated three generative vision-language models on three institutional chest-radiograph corpora and six findings under two elicitation protocols, comprising more than 345,000 finding-level predictions, and estimated finding-by-direction reference agreement at a receiving institution from a small budget of local labels. Estimation strategies were then stress-tested under repeated strict institution-held-out evaluation. Under evaluation excluding the receiving institution from development entirely, adaptive selection among the seven estimators that design admits did not improve on simple fixed alternatives: it achieved a mean Brier score of 0.1083, against 0.0853 for always using a Beta-Binomial empirical-Bayes estimator and 0.0855 for a target-only logistic model. Those two differ by 0.0003, less than this family's own sensitivity to a change of solver version, and each leads in about half the settings, so no default can be recommended. Their advantage over estimators pooling across institutions was concentrated at one site and not confirmatory once clustered by institution, and a plug-in empirical-Bayes posterior-predictive count interval at a nominal 95% level covered 87.0%, less at the hardest institution. Reference agreement therefore has to be re-evaluated per site and per interface; these results concern agreement with institutional labels, not clinical correctness.
Pengyang Yu, Yiou Wang, Zhongping Dong +3
School of Computer Science, University College Dublin, Dublin, Ireland · Department of Medical Imaging, The Third Affiliated Hospital of Southern Medical University, Guangzhou, China · Dublin City University, Dublin, Ireland