cs.CVOct 27, 2025

Accurate and Scalable Multimodal Pathology Retrieval via Attentive Vision-Language Alignment

Authors: Hongyi WangZhengjie ZhuJunlin HouJiabo MaFang WangYue ShiQiuyu CaiJili Wang+8 more

Organizations: Department of Computer Science and Engineering, The Hong Kong University of Science and Technology, Hong Kong SAR, China · College of Computer Science and Technology, Zhejiang University, Hangzhou, China. · School of Medicine, Wake Forest University, Winston-Salem, NC, USA. · Department of Radiology, Union Hospital, Tongji Medical College, Huazhong University of Science and Technology, Wuhan, China · Department of Pathology, Sir Run Run Shaw Hospital, School of Medicine, Zhejiang University, Hangzhou, China · Department of Pathology, The First Affiliated Hospital, School of Medicine, Zhejiang University, Hangzhou, China. · Department of Pathology, The Central Hospital of Wuhan, Tongji Medical College, Huazhong University of Science and Technology, Wuhan, China · Department of Pathology, Nanfang Hospital, School of Medicine, Southern Medical University, Guangzhou, China · College of Information Science and Engineering, Ritsumeikan University, Osaka, Japan. · Zhejiang Key Laboratory of Multi-omics Precision Diagnosis and Treatment of Liver Diseases, Zhejiang University, Hangzhou 310063, China · Department of Chemical and Biological Engineering, The Hong Kong University of Science and Technology, Hong Kong SAR, China · Division of Life Science, The Hong Kong University of Science and Technology, Hong Kong SAR, China · Shenzhen-Hong Kong Collaborative Innovation Research Institute, The Hong Kong University of Science and Technology, Shenzhen, China · State Key Laboratory of Nervous System Disorders, The Hong Kong University of Science and Technology, Hong Kong SAR, China

Abstract

The rapid digitization of histopathology slides has opened new opportunities for computational tools in clinical and research workflows. Content-based slide retrieval can help pathologists identify morphologically and semantically related precedent cases, supporting expert diagnosis and example-based education. Effective retrieval of whole-slide images (WSIs), however, remains challenging because gigapixel slides contain abundant irrelevant content, focal diagnostic patterns and slide-level semantic information that must be represented at a practicable search cost. Here we present PathSearch, a retrieval framework that combines fine-grained attentive mosaics with slide-level embeddings aligned through vision-language contrastive learning. Trained on 6,926 slide-report pairs, PathSearch captures both fine-grained morphological cues and high-level semantic patterns to enable accurate and flexible retrieval. The framework supports two key functionalities: (1) mosaic-based image-to-image (I2I) retrieval, ensuring accurate and efficient slide search; and (2) multimodal retrieval, where text queries can directly retrieve relevant slides. PathSearch was evaluated on eight tasks comprising 5,021 evaluation slides, spanning malignancy assessment on frozen and hematoxylin and eosin (H&E)-stained slides, lymph-node metastasis detection, tumor subtyping, mixed-gallery rare-cancer retrieval, and hepatocellular carcinoma (HCC) risk stratification. Internal and external experimental results demonstrate that PathSearch consistently outperforms the strongest existing methods without compromising multimodal accuracy. A multi-center reader study further demonstrated increases in task-level mean diagnostic accuracy, confidence, and inter-observer agreement with PathSearch's support. Together, these results support the effectiveness of PathSearch across diverse retrieval tasks and evaluation settings.

Explore similar work

CardsList