Efficient Multi-Granularity Knowledge Transfer for Radiology Report Generation
Organizations: Rearch Institute of United Imaging, Shenzhen 518000, China · Global College, Shanghai Jiao Tong University, Shanghai 200240, China · Shenzhen Institutes of Advanced Technology, Chinese Academy of Sciences, Shenzhen 518000, China · Department of Radiology & Institute for Medical Imaging Technology, Ruijin Hospital, Shanghai Jiao Tong University School of Medicine, 197 Ruijin 2nd Road, Huangpu District, Shanghai, China
Abstract
Radiology report generation can automatically generate clinical descriptions from X-ray images, thereby significantly improving the efficiency of radiologists. This task is challenging because it requires medical knowledge to accurately identify diseases and describe them in a professional manner. However, existing methods often overlook the importance of enhancing medical knowledge in describing pivotal areas, a capability that requires models to effectively extract and aggregate knowledge at multiple levels of granularity. Accordingly, we herein propose a novel and compact Efficient Multi-Granularity Knowledge Transfer (\textbf{EMGKT}) method to address the above issues. First, we encode global knowledge embeddings using a medical vision-language model, which provides contextual medical knowledge. Moreover, we devise a novel Fine-Grained Knowledge Distillation (FGKD) training task which efficiently extract fine-grained knowledge. Specifically, the FGKD training task contains teacher embeddings and student embeddings. Teacher embeddings are encoded using extra priors; while student embeddings are learned from the teacher embeddings through knowledge distillation. During inference, the student embeddings are used to enhance fine-grained knowledge while the teacher embeddings are discarded, resulting in negligible computational costs and no need for extra priors. Finally, we further develop a mixture of disease diagnosis expert classifiers to enhance knowledge extraction. The classifiers are initialized using disease embeddings and are modeled as different experts to address various granularity features. Notably, \textbf{EMGKT} can be efficiently applied to most existing methods. Extensive experiments are conducted on two widely-used public datasets and various baselines, which demonstrates the effectiveness and transferability of \textbf{EMGKT}.
Figures & tables
| Method | NLG Metrics | |||||
|---|---|---|---|---|---|---|
| BLEU-1 | BLEU-2 | BLEU-3 | BLEU-4 | METEOR | ROUGE-L | |
| Experimental results on MIMIC-CXR dataset. | ||||||
| R2GenCMN * [ 6 ] | 0.344 | 0.210 | 0.139 | 0.098 | 0.136 | 0.275 |
| +Ours | 0.356 | 0.219 | 0.146 | 0.105 | 0.139 | 0.278 |
| 0.012 | 0.009 | 0.007 | 0.007 | 0.003 | 0.003 | |
| PromptMRG * [ 22 ] | 0.395 | 0.236 | 0.154 | 0.112 | 0.154 | 0.267 |
| Method | Precision | Recall | F1-Score |
|---|---|---|---|
| R2GenCMN [ 6 ] | 0.334 | 0.275 | 0.278 |
| GSKET [ 40 ] | 0.458 | 0.348 | 0.371 |
| Clinical-BERT [ 42 ] | 0.397 | 0.435 | 0.415 |
| KiUT [ 34 ] | 0.371 | 0.318 | 0.321 |
| DCL [ 39 ] | 0.471 | 0.352 | 0.373 |
| METransformer [ 38 ] | 0.364 | 0.309 | 0.311 |
| Model | Year | NLG Metrics | CE Metrics | Avg | |||||||
|---|---|---|---|---|---|---|---|---|---|---|---|
| BLEU-1 | BLEU-2 | BLEU-3 | BLEU-4 | METEOR | ROUGE | Precision | Recall | F1 | |||
| R2Gen | ACL 2020 | 0.353 | 0.218 | 0.145 | 0.103 | 0.142 | 0.277 | 0.333 | 0.273 | 0.276 | 0.236 |
| M2TR | ACL 2021 | 0.378 | 0.232 | 0.154 | 0.107 | 0.145 | 0.272 | 0.240 | 0.428 | 0.308 | 0.252 |
| MKSG | MIA 2022 | 0.363 | 0.228 | 0.156 | 0.115 | - | 0.284 | 0.458 | 0.348 | 0.371 | - |
| M2KT | MIA 2023 | 0.386 | 0.237 | 0.157 | 0.111 | - | 0.274 | 0.420 | 0.339 | 0.352 | - |
| ME | CVPR 2023 | 0.386 | 0.250 | 0.169 | 0.124 | 0.152 | 0.291 | 0.364 | 0.309 | 0.311 | 0.262 |
| Model | Year | NLG Metrics | CE Metrics | Avg | |||||||
|---|---|---|---|---|---|---|---|---|---|---|---|
| BLEU-1 | BLEU-2 | BLEU-3 | BLEU-4 | METEOR | ROUGE | Precision | Recall | F1 | |||
| R2Gen | ACL 2020 | 0.289 | 0.155 | 0.087 | 0.052 | 0.128 | 0.243 | 0.151 | 0.145 | 0.145 | 0.155 |
| M2KT | MIA 2023 | 0.371 | 0.239 | 0.151 | 0.078 | 0.153 | 0.261 | 0.153 | 0.145 | 0.145 | 0.188 |
| DCL | CVPR 2023 | 0.354 | 0.230 | 0.148 | 0.074 | 0.152 | 0.267 | 0.168 | 0.167 | 0.162 | 0.191 |
| RGRG | CVPR 2023 | 0.266 | 0.215 | 0.147 | 0.063 | 0.146 | 0.180 | 0.183 | 0.187 | 0.180 | 0.174 |
| CVT2Dis. | Artif.Intell.Med 2022 | 0.383 | 0.236 | 0.157 | 0.082 | 0.147 | 0.277 | 0.174 | 0.172 | 0.168 | 0.200 |
| Multi-Granularity Knowledge Encoder | Mixture of Disease Diagnosis Experts | BLEU-1 | BLEU-2 | BLEU-3 | BLEU-4 | METEOR | ROUGE-L | Inference ms/img |
|---|---|---|---|---|---|---|---|---|
| 0.395 | 0.236 | 0.154 | 0.112 | 0.154 | 0.267 | 124 | ||
| ✓ | 0.405 | 0.246 | 0.162 | 0.114 | 0.154 | 0.272 | 137 | |
| ✓ | ✓ | 0.409 | 0.250 | 0.166 | 0.117 | 0.157 | 0.273 | 137 |
| Global Embeddings | Fine-Grained Vision Embeddings | Fine-Grained Text Embeddings | BLEU-1 | BLEU-2 | BLEU-3 | BLEU-4 | METEOR | ROUGE-L |
|---|---|---|---|---|---|---|---|---|
| 0.395 | 0.236 | 0.154 | 0.112 | 0.154 | 0.267 | |||
| 0.401 | 0.243 | 0.160 | 0.112 | 0.156 | 0.271 | |||
| 0.406 | 0.247 | 0.162 | 0.114 | 0.154 | 0.272 | |||
| 0.409 | 0.250 | 0.166 | 0.117 | 0.157 | 0.273 |
| CLIP Embedding Initialization | Mixture of Experts | BLEU-1 | BLEU-2 | BLEU-3 | BLEU-4 | METEOR | ROUGE-L |
|---|---|---|---|---|---|---|---|
| 0.395 | 0.236 | 0.154 | 0.112 | 0.154 | 0.267 | ||
| ✓ | 0.400 | 0.245 | 0.161 | 0.113 | 0.152 | 0.272 | |
| ✓ | 0.402 | 0.243 | 0.158 | 0.110 | 0.155 | 0.270 | |
| ✓ | ✓ | 0.409 | 0.250 | 0.166 | 0.117 | 0.157 | 0.273 |