cs.CLApr 22, 2026

Weighting What Matters: Boosting Sample Efficiency in Medical Report Generation via Token Reweighting

Authors: Alexander WeersDaniel RueckertMartin J. Menten

Organizations: TUM School of Computation, Information and Technology, Technical University of Munich, DE · Munich Center for Machine Learning (MCML), DE · Department of Computing, Imperial College London, UK

Abstract

Training vision-language models (VLMs) for medical report generation is often hindered by the scarcity of high-quality annotated data. This work evaluates the use of a weighted loss function to improve data efficiency. Compared to standard cross-entropy loss, which treats all token prediction errors equally, the reweighted loss shifts the focus to semantically salient tokens with outsized clinical importance. In experiments on ophthalmological report generation, we show that this simple method improves efficiency across multiple data scales, achieving similar report quality with up to ten times less training data.

Explore similar work

CardsList
  1. SDR: Set-Distance Rewards for Radiology Report Generation

    May 30, 2026Halil Ibrahim Gulluk, Max Van Puyvelde, Wim Van Criekinge +1Chest RadiographyReport