SentZero: An Enhanced Sentence-Centric Vision-Language Pretraining for Multi-Task Zero-Shot Chest X-Ray Analysis
Organizations: Kim Jaechul Graduate School of AI, Korea Advanced Institute of Science and Technology (KAIST)
Abstract
Vision-language (VL) pretraining using paired chest X-ray (CXR) images and radiology reports has shown strong potential for medical image understanding. However, existing methods often remain dependent on task-specific finetuning because radiology reports are lengthy, clinically dense, and difficult to align with simple zero-shot prompts. Recent sentence-level approaches partially address this limitation using clinical phrases extracted by large language models (LLMs), but they largely overlook the intrinsic characteristics of radiology discourse. In particular, limited positive-pair diversity constrains further gains, while clinically equivalent sentences frequently recur across patients, creating false negatives in contrastive learning. To address these issues, we propose SentZero, an enhanced sentence-centric VL pretraining framework for zero-shot, multi-task CXR analysis. SentZero introduces LLM-based abstract-level sentence structuring and mapping to expand positive-pair diversity, together with an additional loss term to mitigate false negatives. We further introduce sentence-conditioned residual modulation of visual embeddings, enabling visual features to adapt to the semantic characteristics of each input sentence. Across diverse downstream tasks and datasets, SentZero improves zero-shot generalization and outperforms prior multi-task zero-shot methods.
Figures & tables
| Method | Venue | Classification | Grounding | |||||
| (AUROC) | (Pointing Acc.) | |||||||
| OpenI | CXR14 | PadChest | CXD10 | CheXPert | CXD10 | MS-CXR | ||
| GLoRIA | ICCV’21 | 0.589 | 0.610 | 0.565 | 0.645 | 0.750 | 0.367 | - |
| BioViL-T | CVPR’23 | 0.702 | 0.729 | 0.655 | 0.708 | 0.789 | 0.351 | 0.719 |
| MedKLIP | ICCV’23 | 0.759 | 0.726 | 0.629 | 0.713 | 0.879 | 0.481 | 0.407 |
| KAD | Nat. Comm.’23 | 0.807 | 0.789 | 0.750 | 0.735 | 0.905 | 0.391 | - |
| Component | Classification | Grounding | ||||||||
| MF | ALM | TC | FN | OpenI | CXR14 | PadChest | CXD10 | CheXPert | CXD10 | MS-CXR |
| 0.8747 | 0.8140 | 0.8592 | 0.8090 | 0.9080 | 0.6513 | 0.8922 | ||||
| ✓ | 0.8741 | 0.8164 | 0.8559 | 0.8093 | 0.9109 | 0.6655 | 0.9222 | |||
| ✓ | ✓ | 0.8856 | 0.8323 | 0.8635 | 0.8395 | 0.8978 | 0.6811 | 0.9162 | ||
| ✓ | ✓ | 0.8828 | 0.8236 | 0.8652 | 0.8226 | 0.9162 | 0.7032 | 0.8922 | ||
| ✓ | ✓ | 0.8774 | 0.8239 | 0.8607 | 0.8230 | 0.9155 | 0.6556 | 0.8862 | ||
| Method | Classification | Grounding | |||||
| OpenI | CXR14 | PadChest | CXD10 | CheXPert | CXD10 | MS-CXR | |
| None | 0.8858 | 0.8376 | 0.8627 | 0.8431 | 0.9047 | 0.7121 | 0.9042 |
| False Negative Masking | 0.8815 | 0.8334 | 0.8642 | 0.8444 | 0.8981 | 0.6923 | 0.8982 |
| False Negative Transition | 0.8492 | 0.8139 | 0.7904 | 0.7951 | 0.9067 | 0.5138 | 0.7066 |
| Ours | 0.8891 | 0.8389 | 0.8609 | 0.8433 | 0.9037 | 0.7315 | 0.9222 |
Appendix figures & tables6 assets
Supplementary material from the paper’s appendix.
Appendix
| Classification | Grounding | ||||||
| OpenI | CXR14 | PadChest | CXD10 | CheXPert | CXD10 | MS-CXR | |
| w/o text feature | 0.8804 | 0.8358 | 0.8599 | 0.8523 | 0.9000 | 0.7298 | 0.9042 |
| w/ text feature | 0.8891 | 0.8389 | 0.8609 | 0.8433 | 0.9037 | 0.7315 | 0.9222 |
| Classification | Grounding | ||||||
| OpenI | CXR14 | PadChest | CXD10 | CheXPert | CXD10 | MS-CXR | |
| 0.01 | 0.8891 | 0.8389 | 0.8609 | 0.8433 | 0.9037 | 0.7315 | 0.9222 |
| 0.03 | 0.8851 | 0.8381 | 0.8637 | 0.8490 | 0.9074 | 0.7018 | 0.8922 |
| 0.05 | 0.8882 | 0.8377 | 0.8591 | 0.8510 | 0.9047 | 0.7209 | 0.9341 |
| Classification | Grounding | ||||||
| OpenI | CXR14 | PadChest | CXD10 | CheXPert | CXD10 | MS-CXR | |
| 10% | 0.8815 | 0.8350 | 0.8621 | 0.8457 | 0.9039 | 0.6998 | 0.9281 |
| 20% | 0.8891 | 0.8389 | 0.8609 | 0.8433 | 0.9037 | 0.7315 | 0.9222 |
| 30% | 0.8818 | 0.8335 | 0.8613 | 0.8489 | 0.8951 | 0.7168 | 0.9341 |