cs.CVOct 8, 2026

Efficient Multi-Granularity Knowledge Transfer for Radiology Report Generation

Authors: Xubin Zhong, Zheyu Zhang, Wenjian Qin, Ning Wen

Organizations: Rearch Institute of United Imaging, Shenzhen 518000, China · Global College, Shanghai Jiao Tong University, Shanghai 200240, China · Shenzhen Institutes of Advanced Technology, Chinese Academy of Sciences, Shenzhen 518000, China · Department of Radiology & Institute for Medical Imaging Technology, Ruijin Hospital, Shanghai Jiao Tong University School of Medicine, 197 Ruijin 2nd Road, Huangpu District, Shanghai, China

Abstract

Radiology report generation can automatically generate clinical descriptions from X-ray images, thereby significantly improving the efficiency of radiologists. This task is challenging because it requires medical knowledge to accurately identify diseases and describe them in a professional manner. However, existing methods often overlook the importance of enhancing medical knowledge in describing pivotal areas, a capability that requires models to effectively extract and aggregate knowledge at multiple levels of granularity. Accordingly, we herein propose a novel and compact Efficient Multi-Granularity Knowledge Transfer (\textbf{EMGKT}) method to address the above issues. First, we encode global knowledge embeddings using a medical vision-language model, which provides contextual medical knowledge. Moreover, we devise a novel Fine-Grained Knowledge Distillation (FGKD) training task which efficiently extract fine-grained knowledge. Specifically, the FGKD training task contains teacher embeddings and student embeddings. Teacher embeddings are encoded using extra priors; while student embeddings are learned from the teacher embeddings through knowledge distillation. During inference, the student embeddings are used to enhance fine-grained knowledge while the teacher embeddings are discarded, resulting in negligible computational costs and no need for extra priors. Finally, we further develop a mixture of disease diagnosis expert classifiers to enhance knowledge extraction. The classifiers are initialized using disease embeddings and are modeled as different experts to address various granularity features. Notably, \textbf{EMGKT} can be efficiently applied to most existing methods. Extensive experiments are conducted on two widely-used public datasets and various baselines, which demonstrates the effectiveness and transferability of \textbf{EMGKT}.

Figures & tables

Explore similar work

Mar 11, 2024cs.CV

Multimodal Large Language Model driven Radiology Report Generation with Clinical Knowledge Enhancement

Radiology report generation (RRG) has attracted significant attention due to its potential to reduce the workload of radiologists. The performance of current RRG approaches remains unsatisfactory against clinical standards. This paper introduces a novel RRG method, MLLM-RRG, that integrates multimodal large language models (MLLMs) with various types of clinical knowledge to generate accurate and comprehensive chest X-ray reports. Our method first designs a referring anatomical feature extractor that leverages anatomical knowledge to analyze different regions of the chest X-ray image and extract visual features without explicitly detecting regions. Next, based on the MLLM's decoder, we develop a multimodal report generator that leverages multimodal prompts constructed from dedicated visual features and textual instructions to produce the radiology report in an auto-regressive way. Finally, we introduce a disease-oriented clinical classification and alignment scheme in a multi-task learning manner to leverage disease knowledge to better preserve the clinical relevance among the generated reports. Once the model is trained, we also introduce a novel clinical quality reinforcement learning strategy to enhance the MLLM with report knowledge, further refining the tones of the generated reports towards radiologists. Extensive experiments on the MIMIC-CXR and IU X-Ray datasets demonstrate the superiority of our method over the state of the art. Our codes will be available at https://github.com/viscom-tongji/MLLM-RRG.
Feb 4, 2026cs.CL

LEAD: Layer-wise Expert-aligned Decoding for Faithful Radiology Report Generation

Radiology Report Generation aims to produce accurate and coherent diagnostics from medical images. Although large vision-language models improve report fluency and accuracy, they still suffer from hallucinations by generating plausible pathological descriptions that are not supported by the input images. Existing methods primarily rely on external knowledge guidance to facilitate the alignment between generated text and visual information. However, these approaches often ignore the inherent decoding priors and vision-language alignment biases in pretrained models and lack robustness due to reliance on constructed guidance. In this paper, we propose Layer-wise Expert-aligned Decoding, a method that directly intervenes in the internal decoding process of large vision-language models. A pathology-specific expert module is designed to extract discriminative pathological features, which are then injected into each decoder layer through a gated mechanism. This architecture enables the large language model to progressively incorporate expert features during generation through a learned layer-wise gating function, thereby mitigating decoding biases and steering generation toward factual consistency. Experiments on multiple public datasets demonstrate that the proposed method improves clinical accuracy and factual consistency while maintaining competitive report generation quality.
Jun 1, 2026cs.AI

RL-ACRGNet: Reinforcement Learning-Based Chest Radiology Report Generation Network

Medical imaging interpretation is a foundational pillar of modern clinical diagnostics, yet the manual generation of radiology reports remains a time-consuming process prone to interpretation inconsistencies. Within the field of medical AI, automating these descriptions through deep learning promises to streamline clinical workflows and standardise diagnostic output. However, accurate disease detection and precise report generation remain significant challenges due to limitations in capturing fine-grained visual features and ensuring clinical coherence. To address these issues, we propose RL-ACRGNet, an improved encoder-decoder model that integrates a pre-trained DenseNet encoder with a multilevel LSTM decoder within an off-policy reinforcement learning framework. Using a dual-network approach to refine visual-semantic embeddings through a metric-based reward mechanism, we demonstrate that RL-ACRGNet consistently outperforms state-of-the-art baselines on the IU-Xray dataset, achieving quantitative improvements in BLEU-4 (0.47%), METEOR (0.17%) and ROUGE-L (0.518). Furthermore, comprehensive evaluations on the large-scale MIMIC-CXR data set confirm the robust generalisation of the model and its ability to generate high-quality, clinically relevant reports