Unmentioned Checklist Findings Change How Reinforcement Learning Appears to Improve Chest Radiograph Report Checking
Organizations: Department of Computer Science, University of Rochester, Rochester, NY 14627, USA. · Department of Electrical and Computer Engineering, University of Rochester, Rochester, NY 14627, USA. · Department of Imaging Sciences, University of Rochester Medical Center, Rochester, NY 14642, USA. · Department of Biomedical Engineering, University of Rochester, Rochester, NY 14627, USA.
Abstract
Automated checks of radiology reports may rely on AI-generated checklists that leave findings unmentioned. We used reinforcement learning to train a vision-language model to fill in a 12-finding checklist from a chest radiograph without seeing the sentence under test; a separate checking model judged the sentence from the checklist. On held-out patients, a rule-based check and an independent medical checker, neither used in training, measured discrimination gains (Youden index) of 12.6% and 11.8%; only the rule-based check met the prespecified false-alarm criterion. Switching to the training format, which fixes finding order and enters unmentioned findings as absent, raised the training checker's measured gain and lowered the independent checker's, a prespecified comparison that yielded 6.2% (95% interval 2.0% to 10.5%) and, post hoc on held-out patients, 7.7%. Across 8 checking models, acceptance of a label-consistent negative statement about an unmentioned finding ranged from 1.0% to 97.0%. Labels were report-derived, not radiologist-adjudicated.
Figures & tables
| Term | Meaning in this paper |
|---|---|
| Checklist model | The vision-language model that looks at 1 frontal radiograph and fills in the checklist. It never sees the report sentence under test and is the only model that is trained. |
| Checklist | Up to 12 listed findings, each entry giving present or absent, side, an optional box, and a short note. Findings may be left unmentioned. |
| Checker | Any of the 3 methods that give a verdict on a sentence from a checklist: the training checker, the independent checker, and the rule-based check. |
| Checking model | A frozen language model that reads only the checklist and 1 report sentence and answers supported or unsupported. The training checker supplied the reward; the independent checker, MedGemma, never did. |
| Rule-based check | Supports the sentence when the checklist’s present or absent mark for the sentence’s finding agrees with the sentence. It reads the repaired checklist, counting an unmentioned finding as absent, and uses no language model. |
| Report-derived label | Present or absent for each finding, extracted automatically from the report text by the CheXpert labeler, not a radiologist’s reading of the image. |
| Checker | Untrained | Untrained FA | [95%] | FA (upper 95%) | Gain | False alarms |
|---|---|---|---|---|---|---|
| Validation pairs, 237 | ||||||
| Rule-based check | 3.0 | 45.1 | 8.3 [ 2.0, 14.1] | 1.7 ( 6.6) | met | not met |
| Independent checker | 2.1 | 36.3 | 7.0 [ 0.4, 13.3] | 4.4 ( 9.6) | met | not met |
| Training checker | 4.6 | 46.0 | 3.8 [ 1.0, 8.2] | 12.0 ( 15.9) | not in criterion | |
| Held-out test pairs, 273 | ||||||
| Rule-based check | 3.3 | 40.7 | 12.6 [ 6.9, 18.2] | 2.7 ( 1.9) | met | met |
| Training checker | Independent checker | |||
| Format of untrained and trained checklists | FA | FA | ||
| As generated (evaluation default) | 3.8 [ 1.0, 8.2] | 12.0 [ 7.6, 16.5] | 7.0 [ 0.4, 13.3] | 4.4 [ 1.4, 10.5] |
| Fixed order, unmentioned left out | 3.7 [ 1.0, 8.2] | 12.0 [ 7.4, 16.6] | 6.2 [ 0.3, 12.8] | 4.9 [ 0.9, 10.9] |
| Model’s order, absences entered | 5.8 [ 0.8, 10.8] | 6.4 [ 1.4, 11.5] | 5.4 [ 1.0, 12.0] | 0.8 [ 4.5, 6.5] |
| Training format (fixed order, absences entered) | 6.0 [ 1.0, 10.8] | 6.9 [ 1.9, 11.9] | 3.1 [ 3.7, 9.5] | 1.8 [ 3.9, 7.9] |
| Checker difference, training format (prespecified) | 6.2 [ 2.0, 10.5] on 237 pairs | |||
| Negative accepted | Difference from the training checker | ||
|---|---|---|---|
| Checker | (finding unmentioned) | Validation | Held-out |
| Training checker | 64.0 / 53.5 | – | – |
| MedGemma-4B | 79.0 / 74.0 | 6.2 [ 2.0, 10.5] | 7.7 [ 2.8, 12.8] |
| Gemma-3-4B | 1.0 / 1.6 | 0.6 [ 4.0, 4.8] | 4.9 [ 0.2, 10.2] |
| Gemma-3-12B | 91.0 / 91.3 | 1.5 [ 1.6, 4.7] | 3.2 [ 0.1, 6.3] |
| Gemma-3-27B | 56.0 / 50.4 | 0.3 [ 2.4, 3.0] | 1.9 [ 1.1, 5.0] |
Appendix figures & tables17 assets
Supplementary material from the paper’s appendix.
Appendix
| Training judge | Independent judge | Rule-based reader | ||||
|---|---|---|---|---|---|---|
| Record | FS | FS | FS | |||
| Untrained writer | 4.6 | 45.8 | 2.1 | 36.3 | 2.9 | 45.0 |
| Development run, step 40 | 8.8 | 53.8 | 12.2 | 35.9 | 12.6 | 44.1 |
| Change | 4.2 | 8.0 | 10.1 | 0.4 | 9.7 | 0.8 |
| Labels as flags | 68.5 | 27.7 | 53.4 | 4.2 | 100.0 | 0.0 |
| Training judge | Independent judge | |||
| Records read | FS | FS | ||
| Original form | 4.2 [ 1.7, 10.6] | 8.0 [ 4.1, 12.1] | 10.1 [ 2.1, 18.1] | 0.4 [ 5.6, 4.8] |
| Original form, sides not applicable | 3.8 [ 2.4, 10.2] | 7.6 [ 3.7, 11.5] | 11.4 [ 3.4, 19.2] | 0.4 [ 5.5, 4.6] |
| Reward-scoring form, every field | 6.7 [ 0.4, 13.2] | 6.3 [ 2.1, 10.4] | 6.3 [ 1.3, 13.6] | 0.4 [ 5.0, 4.3] |
| Reward-scoring form, notes removed | 7.6 [ 1.7, 13.6] | 3.4 [ 0.8, 7.6] | 5.9 [ 1.7, 13.7] | 2.5 [ 8.0, 2.9] |
| Reward-scoring form, flags only | 8.0 [ 1.7, 14.5] | 2.9 [ 1.7, 7.5] | 3.4 [ 3.0, 10.1] | 3.4 [ 1.7, 8.2] |
| Development run | First registered run | |
|---|---|---|
| Finding | images / groups / informative | images / groups / informative |
| Pleural effusion | 252 / 265 / 118 (45) | 256 / 273 / 124 (45) |
| Edema | 229 / 229 / 110 (48) | 283 / 283 / 120 (42) |
| Consolidation | 200 / 211 / 68 (32) | 200 / 213 / 61 (29) |
| Cardiomegaly | 140 / 140 / 46 (33) | 158 / 158 / 24 (15) |
| Atelectasis | 132 / 132 / 21 (16) | 167 / 167 / 30 (18) |
| Item | Setting |
|---|---|
| Writer and judge | Qwen3-VL-8B-Instruct, bfloat16; judge frozen and served separately |
| Adapter | low-rank adaptation (LoRA) rank 16, , dropout 0, on q, k, v, o, gate, up, down projections of the language model; vision tower untouched |
| Algorithm | group relative policy optimization, Kullback-Leibler (KL) loss to the initial writer (coefficient 0.01, low-variance estimator), no entropy bonus |
| Dynamic sampling | groups with no reward variation discarded and the batch refilled |
| Rollouts | 16 records per image, temperature 1.0, up to 1,024 new tokens |
| Batch | 8 images per step, mini-batch 4, learning rate , 40 steps, adapters saved every 10 steps |
| Validation, 237 pairs | Held-out test, 273 pairs | ||||
|---|---|---|---|---|---|
| Reader | Writer | FS | FS | ||
| Rule-based reader | untrained , FS | 3.0 | 45.1 | 3.3 | 40.7 |
| run 1 | 11.0 [ 3.4, 18.6] | 0.4 | 13.6 [ 7.0, 20.1] | 4.0 | |
| run 2 | 9.3 [ 2.5, 16.0] | 3.0 | 11.0 [ 4.5, 17.4] | 2.6 | |
| run 3 | 4.6 [ 1.7, 11.0] | 1.7 | 13.2 [ 6.1, 20.7] | 1.5 | |
| mean | 8.3 [ 2.0, 14.1] | 1.7 | 12.6 [ 6.9, 18.2] | 2.7 | |
| Sentence (pairs) | Judge | Original form | Not assessed | Absent | Marking | Assertion |
|---|---|---|---|---|---|---|
| Denies the finding (148) | Training | 15.8 | 20.0 | 10.6 | 4.3 [ 6.8, 2.0] | 9.5 [ 6.0, 12.7] |
| Independent | 13.3 | 9.7 | 8.3 | 3.6 [ 0.5, 6.6] | 1.4 [ 1.3, 4.0] | |
| Asserts the finding (88) | Training | 5.7 | 0.8 | 0.8 | 4.9 [ 0.8, 9.5] | 1.5 [ 1.1, 4.9] |
| Independent | 11.0 | 11.4 | 11.7 | 0.4 [ 5.7, 5.4] | 0.4 [ 4.1, 4.5] |
| Training judge | Independent judge | |||
|---|---|---|---|---|
| Form of untrained and trained held-out test records | FS | FS | ||
| Original form | 6.4 [ 2.1, 10.8] | 7.2 [ 2.2, 11.9] | 12.5 [ 6.0, 18.7] | 3.8 [ 2.3, 9.5] |
| Fixed order, omissions left out | 6.7 | 6.7 | 11.0 | 4.7 |
| Writer’s order, omissions absent | 11.2 | 1.9 | 10.5 | 0.5 |
| Reward-scoring form | 11.2 [ 6.8, 16.1] | 2.0 [ 3.2, 6.9] | 9.6 [ 3.7, 15.6] | 1.0 [ 7.0, 4.4] |
| Judge | Pairs | original | reward-scoring | Form effect, | Fill effect, FS | Minus training judge |
|---|---|---|---|---|---|---|
| Validation pairs | ||||||
| Qwen3-VL-8B (training) | 236 | 3.8 | 6.1 | 2.3 [ 5.2, 0.6] | 5.4 [ 2.6, 8.0] | |
| MedGemma-4B (independent) | 236 | 7.2 | 3.2 | 4.0 [ 0.1, 8.0] | 3.3 [ 0.6, 6.2] | 6.2 [ 2.0, 10.6] |
| Qwen3-VL-30B-A3B | 238 | 3.9 | 6.3 | 2.4 [ 5.6, 1.2] | 6.3 [ 3.3, 9.0] | 0.1 [ 2.7, 2.4] |
| Phi-4-mini | 238 | 0.6 | 1.0 | 0.4 [ 4.8, 4.0] | 2.2 [ 0.3, 4.5] | 1.8 [ 2.4, 6.1] |
| OLMo-2-7B | 238 | 4.2 | 5.7 | 1.5 [ 5.8, 2.8] | 5.7 [ 2.5, 8.6] | 0.7 [ 3.5, 4.7] |
| Form effect on | Fill effect on FS | ||||||
| Judge | Pairs | Default | Convention | Change | Default | Convention | Change |
| Validation pairs | |||||||
| Qwen3-VL-8B (training) | 238 | 2.2 | 1.7 | 0.6 [ 1.6, 2.5] | 5.3 | 1.6 | 3.7 [ 5.3, 2.0] |
| Qwen3-VL-30B-A3B | 238 | 2.4 | 2.8 | 0.4 [ 2.6, 1.7] | 6.3 | 6.5 | 0.2 [ 1.2, 2.0] |
| MedGemma-4B (independent) | 234 | 3.6 | 5.7 | 2.1 [ 0.8, 5.5] | 3.3 | 2.5 | 0.8 [ 2.3, 0.6] |
| OLMo-2-7B | 238 | 1.5 | 3.6 | 2.1 [ 5.4, 1.6] | 5.7 | 6.4 | 0.7 [ 1.0, 2.3] |
| Comparison | Population | Pairs | Form effect, judge B | Form effect, judge A | B minus A |
|---|---|---|---|---|---|
| Gemma-3-4B vs training judge | validation | 238 | 1.7 [ 5.8, 2.5] | 2.2 [ 5.0, 0.7] | 0.6 [ 4.0, 4.8] |
| Gemma-3-4B vs training judge | held-out test | 274 | 0.1 [ 4.9, 5.1] | 4.7 [ 8.1, 1.7] | 4.9 [ 0.2, 10.2] |
| Gemma-3-12B vs training judge | validation | 238 | 0.7 [ 4.1, 2.8] | 2.2 [ 5.0, 0.7] | 1.5 [ 1.6, 4.7] |
| Gemma-3-12B vs training judge | held-out test | 274 | 1.6 [ 4.0, 0.6] | 4.7 [ 8.1, 1.7] | 3.2 [ 0.1, 6.3] |
| Gemma-3-27B vs training judge | validation | 238 | 2.0 [ 5.2, 1.3] | 2.2 [ 5.0, 0.7] | 0.3 [ 2.4, 3.0] |
| Gemma-3-27B vs training judge | held-out test | 274 | 2.8 [ 6.1, 0.5] | 4.7 [ 8.1, 1.7] | 1.9 [ 1.1, 5.0] |
| Target finding not mentioned by the record | Target finding listed | ||||||
|---|---|---|---|---|---|---|---|
| Judge | Population | Correct negative | Incorrect negative | Correct positive | Incorrect positive | Correct negative | Incorrect negative |
| Training judge | validation | 64.0 99.0 | 65.8 97.5 | 6.6 6.6 | 4.8 8.3 | 64.1 64.1 | 52.4 52.6 |
| MedGemma-4B | validation | 79.0 99.0 | 71.8 97.4 | 15.8 31.6 | 24.1 39.8 | 67.9 72.6 | 56.3 61.3 |
| Qwen3-VL-30B | validation | 40.0 99.0 | 44.3 96.2 | 6.6 6.6 | 4.8 8.3 | 63.5 63.7 | 52.4 52.6 |
| Phi-4-mini | validation | 97.0 99.0 | 83.5 97.5 | 17.1 14.5 | 17.9 16.7 | 86.7 91.3 | 81.6 88.2 |
| OLMo-2-7B | validation | 45.0 97.0 | 49.4 89.9 | 3.9 18.4 | 3.6 14.3 | 64.9 65.1 | 53.6 54.4 |
| Judge | Form | Verdict and reason |
|---|---|---|
| Training | in the original form | rejected, “The record does not mention pleural effusions, so the sentence cannot be confirmed as true based on available data.” |
| Training | reward-scoring form | accepted, “Pleural Effusion is marked as absent in the record.” |
| Independent | in the original form | rejected, “The record does not mention pleural effusions.” |
| Independent | reward-scoring form | accepted, “The record indicates that Pleural Effusion is not present.” |
| , older / newer writer | Omission form | Encoding | Comprehension | ||
| Judge | blank | absent | absent blank | sparse complete | complete / sparse |
| Qwen3-VL-8B (training judge’s model) | 90.0 / 88.8 | 91.6 / 89.4 | 1.0 [ 2.5, 0.6] | 0.1 [ 1.3, 1.1] | 100 / 100 |
| Gemma-3-12B | 89.7 / 87.4 | 91.8 / 89.1 | 0.3 [ 2.2, 1.6] | 0.1 [ 1.1, 1.3] | 100 / 100 |
| Qwen3-VL-30B-A3B | 88.3 / 88.0 | 91.9 / 89.7 | 1.9 [ 3.3, 0.4] | 0.3 [ 0.7, 1.2] | 100 / 97.9 |
| Phi-4-mini | 85.8 / 85.5 | 63.8 / 69.3 | 5.8 [ 3.0, 8.5] | 6.8 [ 9.7, 3.9] | 87.5 / 100 |
| OLMo-2-7B | 61.6 / 60.5 | 60.1 / 56.5 | 2.5 [ 5.7, 0.6] | 2.8 [ 6.0, 0.3] | 87.5 / 91.7 |