cs.CVApr 30, 2026

Assessing Pancreatic Ductal Adenocarcinoma Vascular Invasion: the PDACVI Benchmark

Authors: M. Riera-Marín, O. K. Sikha, J. Rodríguez-Comas, M. S. May, T. Kirscher, X. Coubez, P. Meyer, S. Faisan, +18 more

Organizations: BCN Medtech, Universitat Pompeu Fabra, Barcelona, Spain · Universitätsklinikum Erlangen, Department of Radiology of the Uniklinikum Erlangen (UKER), Erlangen, Germany · University Hospital Erlangen, Imaging Science Institute, Erlangen, Germany · ICUBE Laboratory, CNRS UMR-7357, University of Strasbourg, Strasbourg, France · CLCC Institut-Strauss, Strasbourg, France · Shenzhen Institute of Advanced Technology, Chinese Academy of Sciences, Shenzhen, China · University of Chinese Academy of Sciences, Beijing, China · Universit´e de Rennes 1, CLCC Eug`ene Marquis, and INSERM UMR 1099 LTSI, Rennes, France · German Cancer Research Center (DKFZ), Medical Image Computing (E230), Heidelberg, Germany · Hospital de Sant Pau i la Santa Creu, Diagnostic Imaging Department, Barcelona, Spain · Institut de Recerca Sant Pau - CERCA, Advanced Medical Imaging, Artificial Intelligence, and Imaging-Guided Therapy Research Group, Barcelona, Spain · Instituci´o Catalana de Recerca i Estudis Avan¸ats (ICREA), Barcelona, Spain · TECNALIA, Basque Research and Technology Alliance (BRTA), Bizkaia, Spain

Abstract

Surgical resection remains the only potentially curative treatment for pancreatic ductal adenocarcinoma (PDAC), and eligibility depends on accurate assessment of vascular invasion (VI), i.e., tumor extension into adjacent critical vessels. Despite its importance for preoperative staging and surgical planning, computational VI assessment remains underexplored. Two major challenges are the lack of public datasets and the diagnostic ambiguity at the tumor-vessel interface, which leads to substantial inter-rater variability even among expert radiologists. To address these limitations, we introduce the CURVAS-PDACVI Dataset and Challenge, an open benchmark for uncertainty-aware AI in PDAC staging based on a densely annotated dataset with five independent expert annotations per scan. We also propose a multi-metric evaluation framework that extends beyond spatial overlap to include probabilistic calibration and VI assessment. Evaluation of six state-of-the-art methods shows that strong global volumetric overlap does not necessarily translate into reliable performance at clinically critical tumor-vessel interfaces. In particular, methods optimized for binary segmentation perform competitively on average overlap metrics, but often degrade in high-complexity cases with low expert consensus, either collapsing in volume or overextending at uncertain boundaries. In contrast, methods that model inter-rater disagreement produce better calibrated probabilistic maps and show greater robustness in these ambiguous cases. The benchmark highlights the limitations of volumetric accuracy as a proxy for localized surgical utility, motivating uncertainty-aware probabilistic models for preoperative decision-making.

Explore similar work

CardsList