cs.CVSep 30, 2026

PLRS-IC: A Dual-Calibration Framework for Chest X-Ray Vision-Language Alignment

Authors: Qixing Zhao, Jinpeng Li

Organizations: South China University of Technology

Abstract

Fine-grained vision-language alignment in chest radiography enables zero-shot classification, grounding, and segmentation without task-specific annotations. However, this alignment is fundamentally hindered by two intertwined sources of ambiguity: projection-induced visual mismatch and patient-agnostic semantic overlap. First, at the local feature level, frontal and lateral radiographs exhibit distinct appearances for the same clinical finding, rendering a shared patch-text similarity geometry inherently suboptimal. Compounding this visual ambiguity is a semantic mismatch during global contrastive optimization, where instance-level objectives penalize cross-patient pairs as strict negatives even when they share identical positive clinical concepts. To address this dual ambiguity, we propose PLRS-IC, a unified dual-calibration framework for chest X-ray representation learning. At the local alignment stage, Projection-Conditioned Low-Rank Residual Similarity (PLRS) dynamically adapts patch-text matching to projection-specific manifolds using a bounded, parameter-efficient low-rank residual. At the global optimization stage, Information-Content-Calibrated Soft False-Negative Suppression (IC-SFNS) leverages a corpus-derived information-theoretic prior to soften the penalty of semantically overlapping negatives without altering original contrastive assignments. Extensive experiments across nine public zero-shot benchmark settings demonstrate that our framework yields consistent improvements in classification, grounding, and segmentation, validating the necessity of dual-calibration in medical vision-language pre-training.

Figures & tables

Explore similar work

CardsList
  1. GLINT: Sparsely Gated Vision-Language Alignment for Fine-Grained Radiology Representations

    Jun 2, 2026Jonggwon Park, Seongeun Lee, Junhyun Park +6Vision-Language AlignmentFew-Shot Segmentation

  2. SentZero: An Enhanced Sentence-Centric Vision-Language Pretraining for Multi-Task Zero-Shot Chest X-Ray Analysis

    Sep 28, 2026Hangyul Yoon, Hyungyung Lee, Edward Choi +1Chest X-RayRecent Vision-Language Models

  3. AlphaRAD: Grounded Zero-Shot Classification in Chest Radiology via αα-Corrected Binary Cross Entropy and Factorized Latent Supervision

    Sep 1, 2026Jianzhong You, Yuan Gao, Chris McIntoshChest X-RayMedical Vision-Language Models