Efficient Multi-Granularity Knowledge Transfer for Radiology Report Generation
Authors: Xubin Zhong, Zheyu Zhang, Wenjian Qin, Ning Wen
Organizations: Rearch Institute of United Imaging, Shenzhen 518000, China · Global College, Shanghai Jiao Tong University, Shanghai 200240, China · Shenzhen Institutes of Advanced Technology, Chinese Academy of Sciences, Shenzhen 518000, China · Department of Radiology & Institute for Medical Imaging Technology, Ruijin Hospital, Shanghai Jiao Tong University School of Medicine, 197 Ruijin 2nd Road, Huangpu District, Shanghai, China
Radiology report generation can automatically generate clinical descriptions from X-ray images, thereby significantly improving the efficiency of radiologists. This task is challenging because it requires medical knowledge to accurately identify diseases and describe them in a professional manner. However, existing methods often overlook the importance of enhancing medical knowledge in describing pivotal areas, a capability that requires models to effectively extract and aggregate knowledge at multiple levels of granularity. Accordingly, we herein propose a novel and compact Efficient Multi-Granularity Knowledge Transfer (\textbf{EMGKT}) method to address the above issues. First, we encode global knowledge embeddings using a medical vision-language model, which provides contextual medical knowledge. Moreover, we devise a novel Fine-Grained Knowledge Distillation (FGKD) training task which efficiently extract fine-grained knowledge. Specifically, the FGKD training task contains teacher embeddings and student embeddings. Teacher embeddings are encoded using extra priors; while student embeddings are learned from the teacher embeddings through knowledge distillation. During inference, the student embeddings are used to enhance fine-grained knowledge while the teacher embeddings are discarded, resulting in negligible computational costs and no need for extra priors. Finally, we further develop a mixture of disease diagnosis expert classifiers to enhance knowledge extraction. The classifiers are initialized using disease embeddings and are modeled as different experts to address various granularity features. Notably, \textbf{EMGKT} can be efficiently applied to most existing methods. Extensive experiments are conducted on two widely-used public datasets and various baselines, which demonstrates the effectiveness and transferability of \textbf{EMGKT}.
Figures & tables
Fig. 1: Comparison of knowledge transfer approaches in radiology report generation. Existing methods achieve knowledge transfer through two primary approaches: (a) general knowledge transfer, which leverages symptom statistics to construct knowledge graph, and (b) image-level knowledge transfer, which utilizes retrieved reports from knowledge database. (c) We propose a methodology that employs extra priors [ 24 ] to parse images and identify discriminative regions during training. This approach efficiently integrates global image context along with knowledge extracted from these critical areas by leveraging a multi-granularity knowledge encoder coupled with a mixture of disease diagnosis expert module.
Fig. 2: Overview of our method EMGKT during training phrase. Multi-granularity Knowledge Encoder (MgKE) efficiently encodes and integrates multi-granularity knowledge embeddings, which includes a novel fine-grained knowledge distillation approach. Mixture of Disease Diagnosis Experts (MoD 2 E) further enhances disease knowledge by initializing classifiers using text embeddings of detailed disease descriptions; and subsequently fuses classification scores using a mixture of experts approach.
Fig. 3: Illustration of disease description generation using LLM.
Method
NLG Metrics
BLEU-1
BLEU-2
BLEU-3
BLEU-4
METEOR
ROUGE-L
Experimental results on MIMIC-CXR dataset.
R2GenCMN * [ 6 ]
0.344
0.210
0.139
0.098
0.136
0.275
+Ours
0.356
0.219
0.146
0.105
0.139
0.278
△
0.012
0.009
0.007
0.007
0.003
0.003
PromptMRG * [ 22 ]
0.395
0.236
0.154
0.112
0.154
0.267
TABLE I: Comparison between baselines and the improved network with EMGKT . △ denotes the improvements compared to the baselines. * denotes our re-implementation of baselines.
Method
Precision
Recall
F1-Score
R2GenCMN [ 6 ]
0.334
0.275
0.278
GSKET [ 40 ]
0.458
0.348
0.371
Clinical-BERT [ 42 ]
0.397
0.435
0.415
KiUT [ 34 ]
0.371
0.318
0.321
DCL [ 39 ]
0.471
0.352
0.373
METransformer [ 38 ]
0.364
0.309
0.311
TABLE II: The comparison of the clinical efficacy metrics on MIMIC-CXR dataset.
Model
Year
NLG Metrics
CE Metrics
Avg
BLEU-1
BLEU-2
BLEU-3
BLEU-4
METEOR
ROUGE
Precision
Recall
F1
R2Gen
ACL 2020
0.353
0.218
0.145
0.103
0.142
0.277
0.333
0.273
0.276
0.236
M2TR
ACL 2021
0.378
0.232
0.154
0.107
0.145
0.272
0.240
0.428
0.308
0.252
MKSG
MIA 2022
0.363
0.228
0.156
0.115
-
0.284
0.458
0.348
0.371
-
M2KT
MIA 2023
0.386
0.237
0.157
0.111
-
0.274
0.420
0.339
0.352
-
ME
CVPR 2023
0.386
0.250
0.169
0.124
0.152
0.291
0.364
0.309
0.311
0.262
TABLE III: Comparison with other SOTA methods on the MIMIC-CXR dataset. The best results are highlighted in bold.
Model
Year
NLG Metrics
CE Metrics
Avg
BLEU-1
BLEU-2
BLEU-3
BLEU-4
METEOR
ROUGE
Precision
Recall
F1
R2Gen
ACL 2020
0.289
0.155
0.087
0.052
0.128
0.243
0.151
0.145
0.145
0.155
M2KT
MIA 2023
0.371
0.239
0.151
0.078
0.153
0.261
0.153
0.145
0.145
0.188
DCL
CVPR 2023
0.354
0.230
0.148
0.074
0.152
0.267
0.168
0.167
0.162
0.191
RGRG
CVPR 2023
0.266
0.215
0.147
0.063
0.146
0.180
0.183
0.187
0.180
0.174
CVT2Dis.
Artif.Intell.Med 2022
0.383
0.236
0.157
0.082
0.147
0.277
0.174
0.172
0.168
0.200
TABLE IV: Comparing the performance of our model with other SOTA methods on the IU X-Ray dataset.
Fig. 4: Overview of EMGKT during inference phrase.
Multi-Granularity Knowledge Encoder
Mixture of Disease Diagnosis Experts
BLEU-1
BLEU-2
BLEU-3
BLEU-4
METEOR
ROUGE-L
Inference ms/img
0.395
0.236
0.154
0.112
0.154
0.267
124
✓
0.405
0.246
0.162
0.114
0.154
0.272
137
✓
✓
0.409
0.250
0.166
0.117
0.157
0.273
137
TABLE V: Ablation experiments of efficient multi-granularity knowledge transfer.
Global Embeddings
Fine-Grained Vision Embeddings
Fine-Grained Text Embeddings
BLEU-1
BLEU-2
BLEU-3
BLEU-4
METEOR
ROUGE-L
0.395
0.236
0.154
0.112
0.154
0.267
✓
0.401
0.243
0.160
0.112
0.156
0.271
✓
✓
0.406
0.247
0.162
0.114
0.154
0.272
✓
✓
✓
0.409
0.250
0.166
0.117
0.157
0.273
TABLE VI: Ablation experiments on the structure of multi-granularity knowledge encoder.
CLIP Embedding Initialization
Mixture of Experts
BLEU-1
BLEU-2
BLEU-3
BLEU-4
METEOR
ROUGE-L
0.395
0.236
0.154
0.112
0.154
0.267
✓
0.400
0.245
0.161
0.113
0.152
0.272
✓
0.402
0.243
0.158
0.110
0.155
0.270
✓
✓
0.409
0.250
0.166
0.117
0.157
0.273
TABLE VII: Ablation experiments on the structure of disease classifiers.
Fig. 5: Visualization results of the baseline and our method. Content rendered in blue font signifies alignment with the ground-truth, whereas content in red font denotes discrepancies or inaccuracies.
Radiology report generation (RRG) has attracted significant attention due to its potential to reduce the workload of radiologists. The performance of current RRG approaches remains unsatisfactory against clinical standards. This paper introduces a novel RRG method, MLLM-RRG, that integrates multimodal large language models (MLLMs) with various types of clinical knowledge to generate accurate and comprehensive chest X-ray reports. Our method first designs a referring anatomical feature extractor that leverages anatomical knowledge to analyze different regions of the chest X-ray image and extract visual features without explicitly detecting regions. Next, based on the MLLM's decoder, we develop a multimodal report generator that leverages multimodal prompts constructed from dedicated visual features and textual instructions to produce the radiology report in an auto-regressive way. Finally, we introduce a disease-oriented clinical classification and alignment scheme in a multi-task learning manner to leverage disease knowledge to better preserve the clinical relevance among the generated reports. Once the model is trained, we also introduce a novel clinical quality reinforcement learning strategy to enhance the MLLM with report knowledge, further refining the tones of the generated reports towards radiologists. Extensive experiments on the MIMIC-CXR and IU X-Ray datasets demonstrate the superiority of our method over the state of the art. Our codes will be available at https://github.com/viscom-tongji/MLLM-RRG.
Miaojing Shi, Tianyu Cen, Zijie Yue +3
College of Electronic and Information Engineering, Tongji University, Shanghai, China
Radiology Report Generation aims to produce accurate and coherent diagnostics from medical images. Although large vision-language models improve report fluency and accuracy, they still suffer from hallucinations by generating plausible pathological descriptions that are not supported by the input images. Existing methods primarily rely on external knowledge guidance to facilitate the alignment between generated text and visual information. However, these approaches often ignore the inherent decoding priors and vision-language alignment biases in pretrained models and lack robustness due to reliance on constructed guidance. In this paper, we propose Layer-wise Expert-aligned Decoding, a method that directly intervenes in the internal decoding process of large vision-language models. A pathology-specific expert module is designed to extract discriminative pathological features, which are then injected into each decoder layer through a gated mechanism. This architecture enables the large language model to progressively incorporate expert features during generation through a learned layer-wise gating function, thereby mitigating decoding biases and steering generation toward factual consistency. Experiments on multiple public datasets demonstrate that the proposed method improves clinical accuracy and factual consistency while maintaining competitive report generation quality.
Ruixiao Yang, Yuanhe Tian, Di Dong +3
Beijing Institute of Technology · Zhongguancun Academy · Zhongguancun Institute of Artificial Intelligence +1
Medical imaging interpretation is a foundational pillar of modern clinical diagnostics, yet the manual generation of radiology reports remains a time-consuming process prone to interpretation inconsistencies. Within the field of medical AI, automating these descriptions through deep learning promises to streamline clinical workflows and standardise diagnostic output. However, accurate disease detection and precise report generation remain significant challenges due to limitations in capturing fine-grained visual features and ensuring clinical coherence. To address these issues, we propose RL-ACRGNet, an improved encoder-decoder model that integrates a pre-trained DenseNet encoder with a multilevel LSTM decoder within an off-policy reinforcement learning framework. Using a dual-network approach to refine visual-semantic embeddings through a metric-based reward mechanism, we demonstrate that RL-ACRGNet consistently outperforms state-of-the-art baselines on the IU-Xray dataset, achieving quantitative improvements in BLEU-4 (0.47%), METEOR (0.17%) and ROUGE-L (0.518). Furthermore, comprehensive evaluations on the large-scale MIMIC-CXR data set confirm the robust generalisation of the model and its ability to generate high-quality, clinically relevant reports
Yogesh Kumar Meena, Saurabh Agarwal, K. V. Arya
Human-AI Interaction (HAIx) Lab, Indian Institute of Technology Gandhinagar, India · Department of Computer Science and Engineering, Madhav Institute of Technology and Science Deemed University (MITS-DU), Gwalior, India · Multimedia and Information Security Research Group, Department of Computer Science and Engineering, ABV-Indian Institute of Information Technology and Management, Gwalior 474015, India