LEAD: Layer-wise Expert-aligned Decoding for Faithful Radiology Report Generation
Organizations: Beijing Institute of Technology · Zhongguancun Academy · Zhongguancun Institute of Artificial Intelligence · University of Science and Technology of China
Abstract
Radiology Report Generation aims to produce accurate and coherent diagnostics from medical images. Although large vision-language models improve report fluency and accuracy, they still suffer from hallucinations by generating plausible pathological descriptions that are not supported by the input images. Existing methods primarily rely on external knowledge guidance to facilitate the alignment between generated text and visual information. However, these approaches often ignore the inherent decoding priors and vision-language alignment biases in pretrained models and lack robustness due to reliance on constructed guidance. In this paper, we propose Layer-wise Expert-aligned Decoding, a method that directly intervenes in the internal decoding process of large vision-language models. A pathology-specific expert module is designed to extract discriminative pathological features, which are then injected into each decoder layer through a gated mechanism. This architecture enables the large language model to progressively incorporate expert features during generation through a learned layer-wise gating function, thereby mitigating decoding biases and steering generation toward factual consistency. Experiments on multiple public datasets demonstrate that the proposed method improves clinical accuracy and factual consistency while maintaining competitive report generation quality.
Figures & tables
| Dataset | Training | Validation | Test |
|---|---|---|---|
| CheXpert Plus | 40,463 | 5,780 | 11,562 |
| MIMIC-CXR | 270,790 | 2,130 | 3,858 |
| IU X-Ray | 2,069 | 296 | 590 |
| Method | Publication/Setting | Decoder | R-L | M | C | F1 | GREEN |
| R2Gen † [ 3 ] | EMNLP20 | Transformer | 0.246 | 0.113 | 0.077 | 0.181 | 0.168 |
| R2GenCMN † [ 30 ] | ACL21 | Transformer | 0.256 | 0.127 | 0.102 | 0.231 | 0.197 |
| R2GenRL † [ 31 ] | ACL22 | Transformer | 0.186 | 0.101 | 0.012 | 0.196 | – |
| XProNet † [ 32 ] | ECCV22 | Transformer | 0.265 | 0.146 | 0.121 | 0.259 | – |
| CvT2DistilGPT2 † [ 34 ] | AIM23 | GPT2 | 0.238 | 0.118 | 0.101 | 0.246 | 0.185 |
| R2GenGPT-Llama2 [ 6 ] | Meta-Rad.23 | Llama2-7B | 0.239 | 0.132 | 0.122 | 0.216 | 0.217 |
| Method | Publication/Setting | M | R-L | C | F1 |
|---|---|---|---|---|---|
| R2Gen [ 3 ] | EMNLP20 | 0.145 | 0.274 | 0.147 | 0.201 |
| CvT2DistilGPT2 [ 34 ] | AIM23 | 0.139 | 0.265 | 0.128 | 0.180 |
| R2GenGPT-Llama2 [ 6 ] | Meta-Rad.23 | 0.151 | 0.276 | 0.145 | 0.249 |
| R2GenGPT-Llama3 [ 6 ] | Meta-Rad.23 | 0.153 | 0.269 | 0.158 | 0.252 |
| Token-Mixer [ 33 ] | IEEE TMI24 | 0.148 | 0.275 | 0.161 | 0.243 |
| SILC [ 39 ] | IEEE TMI24 | 0.150 | 0.277 | 0.163 | 0.232 |
| Dataset | Scale | Method | R-L | M | C | F1 |
|---|---|---|---|---|---|---|
| CheXpert Plus | 2B | Baseline | 0.249 | 0.138 | 0.105 | 0.236 |
| LEAD | 0.258 | 0.140 | 0.138 | 0.242 | ||
| 4B | Baseline | 0.256 | 0.146 | 0.142 | 0.248 | |
| LEAD | 0.262 | 0.148 | 0.153 | 0.260 | ||
| 8B | Baseline | 0.259 | 0.147 | 0.159 | 0.274 | |
| LEAD | 0.275 ∗ | 0.154 ∗ | 0.199 ∗ | 0.346 ∗ |
| Dataset | Backbone | Method | R-L | M | C | F1 |
|---|---|---|---|---|---|---|
| CheXpert Plus | Qwen3-VL [ 24 ] | Baseline | 0.259 | 0.147 | 0.159 | 0.274 |
| LEAD | 0.275 | 0.154 | 0.199 | 0.346 | ||
| Ministral3 [ 46 ] | Baseline | 0.255 | 0.143 | 0.128 | 0.223 | |
| LEAD | 0.261 | 0.143 | 0.138 | 0.283 | ||
| InternVL3.5 [ 47 ] | Baseline | 0.269 | 0.149 | 0.171 | 0.267 | |
| LEAD | 0.272 | 0.150 | 0.179 | 0.306 |
| Configuration | Metrics | |||||
|---|---|---|---|---|---|---|
| Exp. | Proj. | Fuse. | R-L | M | C | F1 |
| - | - | - | 0.259 | 0.147 | 0.159 | 0.274 |
| ✓ | - | - | 0.265 | 0.141 | 0.160 | 0.275 |
| ✓ | Shared | Gate | 0.265 | 0.147 | 0.168 | 0.281 |
| ✓ | Layer | Add | 0.270 | 0.146 | 0.174 | 0.279 |
| ✓ | Prompt | Input | 0.272 | 0.147 | 0.182 | 0.291 |
| Dataset | maAUC | miAUC | ACC |
|---|---|---|---|
| CheXpert Plus | 0.779 | 0.893 | 0.879 |
| MIMIC-CXR | 0.761 | 0.860 | 0.865 |