cs.CVOct 8, 2026

PathLang: A Language-Centered Benchmark for Vision-Language Models in Computational Pathology

Authors: Fanqi Cheng, Kuo Gong, Shangke Liu, Beidi Zhao, Junchao Zhu, Zheyu Zhu, Leiyue Zhao, Fengbei Liu, +10 more

Organizations: New York University, New York, NY, USA · Columbia University, New York, NY, USA · Weill Cornell Medicine, New York, NY, USA · Cornell University, Ithaca, NY, USA · University of British Columbia, Vancouver, BC, Canada · Vanderbilt University, Nashville, TN, USA · University of Pennsylvania, Philadelphia, PA, USA · Johns Hopkins University, Baltimore, MD, USA · New York Medical College, New York, NY, USA · BC Cancer Agency, Vancouver, BC, Canada · The University of Texas MD Anderson Cancer Center, Houston, TX, USA · Cornell Tech, New York, NY, USA

Abstract

Pathology vision-language models (VLMs) have shown strong visual perception ability, but their robustness in the language domain remains poorly characterized. Existing pathology VLM benchmarks largely rely on canonical closed-set prompts or perturb only generic templates, treating language as a fixed evaluation component rather than a variable axis of model behavior. In clinical practice, however, diagnostic language varies across reports, institutions, and candidate diagnoses. We introduce PathLang, a language-centered and clinically grounded zero-shot benchmark. PathLang holds the underlying slides, ground-truth labels, and image-text evaluation direction fixed while systematically varying only the diagnostic language, so that performance differences reflect how a diagnosis is phrased rather than what is imaged. The language variation follows how pathologists actually rephrase diagnoses (terminology, specificity, and reporting style), and all prompts and candidate pools are validated by six board-certified pathologists. PathLang covers four task families: (1) zero-shot classification with image-text alignment analysis, (2) cross-modal retrieval, (3) paraphrase robustness, including semantic-equivalence paraphrases, length and reporting-style variation, and prompt ensembling, and (4) open-vocabulary diagnosis retrieval over four candidate pools with distinct forms of semantic competition. Across nine VLMs and five public datasets spanning four organs, we find that performance is highly sensitive to clinically equivalent paraphrases, varies substantially across forms of semantic competition, and that image-text alignment quality does not necessarily translate into inter-class separability. We release the prompt corpus, candidate pools, pre-computed text embeddings, and evaluation code at https://anonymous.4open.science/r/PathLang.

Figures & tables

Appendix figures & tables42 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Do Pathology Vision-Language Models Truly See Pathology?

    Jul 23, 2026Chengyang Zhang, Wenchuan Zhang, Bo Li +10VLM EvaluationComputational Pathology

  2. PathView-Bench: Can Multimodal Large Language Models Achieve Fine-grained Multiscale Understanding of Pathology Images?

    Jul 30, 2026Zongyi Chen, Yu Liang, Jie Lin +1Computational PathologyMedical VQA