cs.CLSep 30, 2026

Comparison of techniques for fine-tuning open-weight models for entity extraction from radiology reports

Authors: Aawez Mansuri, Kush Mehta, Mohammadreza Chavoshi, Jahanzaib Malik, Theodorus Dapamede, Frank Li, Rohan Isaac, Beatrice Brown-Mulry, +5 more

Organizations: Department of Radiology and Imaging Sciences, Emory University School of Medicine, Atlanta, GA, USA · Department of Computer Science, Emory University, Atlanta, GA, USA

Abstract

Converting free-text radiology reports into structured labels supports cohort building, quality assurance, and monitoring of clinical imaging models, but the strongest label extractors are hosted proprietary models whose use raises privacy, cost, and reproducibility concerns. We asked whether a fine-tuned open-weight model (Gemma-3-12B) can match GPT-4o at multi-label intracranial hemorrhage (ICH) acuity extraction from non-contrast head-CT reports, and which ingredients matter. Using a 2x2 design, we crossed two adaptation strategies (a discriminative classification head, CH; generative instruction fine-tuning, IFT) with two training-data sources (distillation of real GPT-4o-labeled reports; synthetic reports generated by GPT-4o from real exemplars), across five training sizes, benchmarked on 100 expert-adjudicated reports against GPT-4o and the un-tuned open-weight base. The distilled instruction-tuned model (DIFT) matched GPT-4o (macro-F1 0.845 vs 0.850; p = 1.000) and exceeded the base model by 0.178. The decisive factor was the training-data source, not the fine-tuning method: both synthetic-data models failed to exceed the un-tuned open-weight base at any training size and underperformed the distilled models across all acuity classes. Fine-tuning and inference fit within the memory envelope of a single 24 GB consumer GPU. For narrow, high-value clinical label-extraction tasks, distilling real reports, rather than generating synthetic ones, is what closes the gap to a hosted model, enabling a private, low-cost, version-stable on-premises alternative.

Explore similar work

CardsList
  1. Automatic Extraction of Structured Information from Brain MRI Reports Using an Open-Weight Large Language Model

    Jun 5, 2026Kaouther Mouheb, Amos Pomp, Antoine Manenti +9Radiology Report GenerationMedical Vision-Language Models

  2. PromptRad: Knowledge-Enhanced Multi-Label Prompt-Tuning for Low-Resource Radiology Report Labeling

    May 19, 2026Ying-Jia Lin, Tzu-Chin Lo, Ping-Chien Li +3Radiology Report GenerationMulti-Label Classification

  3. Discrete Diffusion Language Models for Interactive Radiology Report Drafting

    Jul 1, 2026Max Van Puyvelde, Halil Ibrahim Gulluk, Wim Van Criekinge +1Radiology Report GenerationDiffusion Language Models