cs.CVSep 1, 2026

AlphaRAD: Grounded Zero-Shot Classification in Chest Radiology via αα-Corrected Binary Cross Entropy and Factorized Latent Supervision

Authors: Jianzhong YouYuan GaoChris McIntosh

Organizations: Peter Munk Cardiac Centre, University Health Network (UHN), Toronto, Canada · Ted Rogers Centre for Heart Research, Toronto, Canada · Vector Institute, Toronto, Canada · Toronto General Hospital Research Institute, UHN, Toronto, Canada · ⋆Work completed during Ph.D. studies at the University of Toronto. · Department of Computer Science, University of Toronto (U of T), Toronto, Canada · Department of Medical Biophysics, U of T, Toronto, Canada · Department of Medical Imaging, U of T, Toronto, Canada

Abstract

Vision-Language Pretrained Models (VLPMs) offer a scalable path to open-vocabulary chest radiology understanding, yet two aspects remain underexplored: how structured clinical semantics extracted from medical reports can reduce in-batch noise during contrastive learning, and how cross-modal fusion can be designed to produce more faithful spatial grounding without added complexity. We introduce AlphaRAD, addressing these opportunities through two contributions. First, we construct a large-scale structured medical concept space from medical reports parsed by a Large Language Model for training, thereby mitigating in-batch learning noise and removing heuristic pair matching in contrastive learning, and thus naturally positioning AlphaRAD as a medical concept discriminator trained via αα-Corrected Binary Cross-Entropy. Second, we propose FLaS (Factorized Latent Supervision), an extremely simple yet effective cross-modal feature fusion module that factorizes VLPM representations into independent subspaces, using dedicated alignment supervision to enhance the expressiveness of spatial grounding without introducing additional model parameters. Through extensive empirical validation, AlphaRAD shows strong zero-shot generalization across diverse chest radiology tasks. Notably, it establishes state-of-the-art average performance across 16 classification benchmarks, while achieving individual state-of-the-art results via distinct gains on 7 grounding/phrase grounding and 3 segmentation datasets.

Explore similar work

CardsList