See, Measure, and Reason: Learning Visually Grounded Reasoning in Pathology
Authors: Chengyang Zhang, Wenchuan Zhang, Bo Li, Mengran Li, Xinyu Liu, Jiaming Yang, Jie Chen, Zhang Zhang, +3 more
Organizations: College of Computer Science, Sichuan University · Department of Pathology and Institute of Clinical Pathology, West China Hospital, Sichuan University · Department of Computer Science, School of Computing, National University of Singapore · School of Intelligent Systems Engineering, Sun Yat-sen University
Pathological assessment relies on recognizing fine-grained visual details in histological images. Vision-language models (VLMs) increasingly support pathology interpretation, yet their ability to perceive these details remains inadequate. This weakness leads to inaccurate cellular observations that can persist even when final answers are correct. In this paper, we propose ASPECT to improve visually grounded reasoning through explicit supervision of cellular appearance and abundance. ASPECT trains intermediate visual tokens through pathology feature reconstruction, cell feature alignment, and count supervision. Three-stage supervised fine-tuning teaches the model to perceive, generate visual tokens, and reason, followed by reinforcement learning that rewards answer correctness and consistency with reported measurements. We also introduce PathoVernier, a benchmark of 759 expert-reviewed questions from five pathology datasets covering four cellular composition tasks. It evaluates both final answers and intermediate measurements to expose errors hidden by answer accuracy. On PathoVernier, ASPECT achieves relative accuracy gains of approximately 19.2% over the strongest baseline, Gemini-3.1-Pro, and 99.3% over its Qwen3-VL-8B backbone, while reducing RAWR, which measures counting errors within correct responses, by 28.1% and 42.7%, respectively. ASPECT also improves over its backbone on three external pathology benchmarks covering classification and question answering beyond cellular composition tasks.
Figures & tables
Figure 1: Correct answers can conceal inaccurate observations. GPT-5.5 and Patho-R1 reverse regional nuclear densities in (a) and answer correctly despite incorrect counts in (b), where Patho-R1 also misapplies the ratio rule. Green contours mark reference nucleus annotations.
Figure 2: Overview of ASPECT. Pathology feature reconstruction, cell feature alignment, and count supervision train the visual-token representations. Three-stage SFT teaches the model to perceive, generate, and reason with explicit measurements. Subsequent RL rewards answer correctness and consistency with the reported counts.
Figure 3: PathoVernier overview. (a) Distribution of 759 questions across organs, sources, magnifications, and tasks; flow widths represent question counts. (b) Example questions with reference counts.
Models
PathoVernier
PathCLS
PathVQA
Quilt-VQA
Acc ↑
CA ↑
RAWR ↓
Acc ↑
Acc ↑
Acc ↑
Closed-source Models
GPT-5.5 ( OpenAI, 2026a )
0.59
0.65
0.76
0.55
0.69
0.69
Gemini-3.1-Pro ( Google DeepMind, 2026 )
0.63
0.71
0.68
0.67
0.78
0.65
Open-source Models
Qwen3-VL-8B ( Bai et al., 2025 )
0.37
0.53
0.85
0.37
0.67
0.58
Table 1: Comparison on PathoVernier and three external pathology benchmarks. ASPECT leads on all PathoVernier metrics and improves over Qwen3-VL-8B on all external benchmarks. Bold and underlined values denote the best and second-best results, respectively. Dashes indicate CA coverage below 30%; Table C.11 reports coverage and unconditional Count Acc.
Figure 4: Qualitative comparison on PathoVernier (left) and PathCLS (right). Responses are condensed and reformatted to highlight key observations and predictions. ASPECT reports more accurate regional counts on PathoVernier and predicts the reference class on PathCLS.
Configuration
Acc ↑
CA ↑
RAWR ↓
Count Acc ↑
ASPECT-SFT
0.70
0.80
0.53
0.44
w/o visual supervision
0.64
0.77
0.59
0.32
w/o pathology feature reconstruction
0.68
0.75
0.56
0.27
w/o cell feature alignment
0.67
0.77
0.54
0.29
w/o count supervision
0.72
0.78
0.57
0.36
Direct Reason training (Stage 3)
0.62
0.74
0.56
0.31
Table 2: Ablations of visual supervision and the SFT curriculum on PathoVernier. All variants are evaluated before RL with matched total update budgets.
Configuration
RL Reward
Acc ↑
CA ↑
RAWR ↓
Count Acc ↑
Consistency ↑
ASPECT-SFT
N/A
0.70
0.80
0.53
0.44
0.94
Answer only
Acc
0.72
0.79
0.55
0.43
0.92
Consistency only
0.5 Con
0.69
0.76
0.51
0.41
0.98
Answer + count score
Acc + 0.5 S
0.74
0.80
0.50
0.45
0.96
ASPECT-8B
Acc + 0.5 Con
0.75
0.81
0.49
0.47
0.99
Table 3: RL reward ablations on PathoVernier. The listed rewards are the internal term r ; all RL variants retain the structural validity gate and the invalid-response penalty.
Appendix figures & tables18 assets
Supplementary material from the paper’s appendix.
Appendix
Source
SFT
RL
Validation
PathoVernier
Visual
VQA
Lizard ( Graham et al., 2021 )
2,520
2,944
559
40
400
PUMA ( Schuiveling et al., 2025 )
2,745
1,010
219
5
142
PanNuke ( Gamper et al., 2020 )
6,780
2,565
567
20
139
CoNSeP ( Graham et al., 2019 )
196
133
27
2
44
NuCLS ( Amgad et al., 2022 )
854
77
0
26
34
Appendix
Table A.1: Data sources and sample counts. SFT examples are separated into visual-supervision and question-answer examples. RL, validation, and PathoVernier columns report question counts. Multiple examples may share an image.
Source
Original label(s)
Target category/categories
Lizard
Epithelial
Tumor; normal epithelial ∗
Connective tissue
Stromal-like
Lymphocyte
Lymphocyte
Plasma, neutrophil, eosinophil
Other inflammatory
PUMA
Tumor
Tumor
Epithelium
Normal epithelial
Appendix
Table A.2: Source-label mappings for cellular supervision. Starred entries retain aggregate labels: feature alignment uses both listed targets, whereas count supervision uses their sum. NuCLS uses these mappings for count supervision only.
Table A.3: Cellular supervision for sources without nucleus-level reference annotations. CRC cell types inherit patch-level tissue labels; PathVQA cell types come from native nucleus predictions.
Skill
Questions
Skill
Questions
Multi-step composition
388
Count comparison
85
Region selection
407
Region comparison
60
Dominant cell type
208
Cell-type comparison
60
Count band prediction
124
Yes/no
40
Total
1,372
Appendix
Table A.4: Skill composition of the RL training collection. Question selection is based on skill-level SFT performance.
Source
Region selection
Region comparison
Cell-type comparison
Multi-step composition
Total
Lizard
82
85
135
98
400
PUMA
45
47
10
40
142
PanNuke
35
36
33
35
139
CoNSeP
17
7
13
7
44
NuCLS
11
11
2
10
34
Total
190
186
193
190
759
Appendix
Table A.5: PathoVernier question counts by source dataset and task.
Organ
Questions
Organ
Questions
Colon
480
Kidney
4
Skin
142
Ovary
3
Breast
60
Liver
2
Bile duct
15
Pancreas
2
Head and neck
11
Prostate
2
Uterus
10
Stomach
2
Appendix
Table A.6: PathoVernier question counts across 18 harmonized organ categories. Colon includes the Colorectal label from CoNSeP.
Source
Resolution ( μ m/pixel)
Patch width ( μ m)
Approx. magnification
Questions
Lizard
0.50
128.0
20×
400
PanNuke, CoNSeP
0.25
64.0
40×
183
PUMA
0.22
56.3
40×
142
NuCLS
0.20
51.2
40×
34
Total
759
Appendix
Table A.7: Image scale by source. Patch width is calculated from the 256×256 -pixel inputs and source resolution. Magnifications are approximate objective equivalents.
Setting
Value
Model and adaptation
Backbone
Qwen3-VL-8B
Pathology feature tokens
8
Cell tokens
6
Visual teachers
UNI; CellViT-SAM-H
Teacher updates
Frozen; targets extracted offline
Appendix
Table B.8: Model and training configurations of ASPECT. SFT stage lengths are reported as updates within each stage.
Benchmark
Evaluation subset
Questions
Local token limit
API token limit
PathoVernier
Full benchmark
759
512
512
PathCLS
Official test
1,632
512
512
PathVQA
H&E, closed-ended test
1,306
512
512
Quilt-VQA
H&E, closed-ended test
314
512
512
Appendix
Table C.9: Evaluation subsets and output token limits for local models and closed-source APIs.
Partition
Stromal-like
Lymphocyte
Epithelial-like
Inflammatory
Quadrants
(0.50,0.85)
(0.35,0.55)
(0.60,0.90)
(0.55,0.80)
Horizontal bands
(0.40,0.80)
(0.30,0.50)
(0.55,0.80)
(0.55,0.75)
Vertical bands
(0.40,0.75)
(0.30,0.50)
(0.60,0.85)
(0.55,0.75)
Appendix
Table C.10: Frozen proportion thresholds (t1,t2) for multi-step composition, estimated separately for each spatial partition and cell type.
Model
CA coverage
Complete-count coverage
NRAWR
Count Acc ↑
GPT-5.5
748/759
749/759
447
0.210
Gemini-3.1-Pro
733/759
750/759
472
0.296
Qwen3-VL-8B
233/759
261/759
106
0.068
Qwen3-VL-32B
682/759
700/759
333
0.178
InternVL3-8B
448/759
449/759
166
0.071
InternVL3-38B
750/759
753/759
340
0.153
Appendix
Table C.11: Evaluation coverage and unconditional count accuracy on PathoVernier. Both coverage measures report counts out of all 759 questions. NRAWR is the number of correct responses with complete reported counts. Count Acc averages S(C,Q) over all questions, including failures and incomplete responses.
Model
Acc ↑
CA ↑
RAWR ↓
Qwen3-VL-8B
0.374 [0.341, 0.408]
0.532 [0.471, 0.595]
0.854 [0.801, 0.901]
Gemini-3.1-Pro
0.626 [0.588, 0.663]
0.715 [0.681, 0.749]
0.680 [0.651, 0.708]
GPT-5.5
0.594 [0.558, 0.631]
0.652 [0.616, 0.689]
0.756 [0.728, 0.784]
ASPECT-8B
0.746 [0.714, 0.776]
0.807 [0.777, 0.836]
0.489 [0.460, 0.518]
Comparison
Acc gain
CA gain
RAWR reduction
ASPECT vs Qwen3-VL-8B
+0.372 [+0.327, +0.415]
+0.275 [+0.207, +0.342]
+0.365 [+0.305, +0.420]
Appendix
Table E.12: Model scores and paired improvements on PathoVernier with 95% bootstrap confidence intervals. Improvements are ASPECT minus baseline for Acc and CA, and baseline minus ASPECT for RAWR; positive values favor ASPECT.
Figure G.4: Complete responses for the PathoVernier multi-step example. Gemini reaches the correct answer with inaccurate counts, whereas ASPECT recovers the reference counts of 7 lymphocytes among 29 nuclei in the selected region.
Figure G.5: Complete responses for the PathCLS example. ASPECT selects the reference class, squamous cell carcinoma (N), while the three comparison models select other classes.
Figure G.6: Regional comparison of tumor nuclei. ASPECT selects the correct ratio band; the baseline responses illustrate errors in count estimation and in translating reported counts into the final answer.
Figure G.7: Whole-image comparison of stromal and inflammatory nuclei. ASPECT correctly selects comparable , while Gemini and Qwen3-VL-8B reach opposite incorrect conclusions from their count estimates.
Figure G.8: Stromal-density selection across four equal-area horizontal bands. ASPECT identifies the reference region, r1 , while Gemini and Qwen3-VL-8B select r2 and r4 , respectively.
Figure G.9: ASPECT failure cases. Left: correct region selection followed by an incorrect proportion band due to denominator overestimation. Right: a near-correct whole-image count accompanies an incorrect regional distribution and region choice. Both final answers are consistent with the reported counts.
Multimodal large language models (MLLMs) are increasingly used to analyze pathology images. However, dominant multimodal benchmarks in pathology mainly score final diagnostic answers, captions, or reports. These evaluations provide limited insight into whether a model understands the multiscale visual content needed for pathology reasoning and decision-making. We introduce PathVU, a vision-anchored benchmark for fine-grained and multiscale visual understanding in computational pathology. Built from 23 public pathology imaging datasets with human-supervised labels and spatial annotations, PathVU evaluates MLLM understanding in two fields of view: Region FOV for high-resolution local regions and Slide FOV for macro whole-slide views. By converting raw annotations into deterministic task targets, PathVU enables programmatic scoring of region localization, visual recognition, quantity estimation, spatial reasoning, and insufficient-context judgment. The benchmark contains 14 VQA-style tasks, 61,673 images, and 308,070 samples across 28 organs and 7,253,526 annotations. Evaluating 18 representative general-purpose, medical-domain, and pathology-oriented MLLMs, we observe substantial limitations even in advanced models on fine-grained visual tasks across multiscale pathology images. PathVU provides a reproducible basis for developing and evaluating pathology MLLMs with explicit multiscale visual understanding.
Zongyi Chen, Yu Liang, Jie Lin +1
National Institute for Data Science in Health and Medicine, Xiamen University · Institute of Artificial Intelligence Xiamen University, Xiamen, China · Department of Computer Science School of Informatics Xiamen University, Xiamen, China
Pathological images are inherently multi-scale, requiring pathologists to integrate evidence from global tissue architecture at low magnification to cellular morphology at higher magnification for accurate diagnosis. While existing pathological datasets for vision-language models (VLMs) include various scales, they often lack explicit cross-scale reasoning objectives. This limitation prevents VLMs from capturing essential cross-scale representations and learning evidence-based reasoning. To bridge this gap, we introduce the first cross-scale training and evaluation paradigm that formulates pathology interpretation as multi-magnification reasoning. However, creating such a task reveals a critical challenge: multi-image visual question answering (VQA) is prone to text-only shortcuts, which allow models to guess answers using magnification-dependent artifacts rather than visual evidence. To address this, we propose a leakage-aware curation pipeline that combines adversarial text-only screening with constraint-guided question design. Using this pipeline, we construct Scale-VQA, a high-quality benchmark with 4,685 multiple-choice questions grounded in 2,537 pathology images across multiple magnification levels. Finally, we present ScaleReasoner-R1, a model trained via reinforcement learning to optimize performance on cross-scale VQA tasks. ScaleReasoner-R1 achieves state-of-the-art performance on our cross-scale reasoning benchmark and generalizes to SOTA performance on established single-scale benchmarks. Findings suggest that even the limited cross-scale supervision can significantly improve pathological understanding. Code is available at https://github.com/iMVR-PL/ScaleReasoner-R1.
Chi Phan, Tianyi Zhang, Qiaochu Xue +5
Department of Electrical and Computer Engineering, National University of Singapore, Singapore · PuzzleLogic Pte Ltd, Singapore · Department of Pathology, Fujian Medical University Cancer Hospital & Fujian Cancer Hospital, Fuzhou, China
Pathological diagnosis is inherently multi-scale, requiring the integration of global tissue architecture at low magnification with cellular morphology at higher magnification. However, existing pathology benchmarks and vision-language models (VLMs) are still largely developed under single-scale settings, limiting their ability to learn clinically meaningful multi-magnification reasoning. Moreover, naively constructed visual question answering (VQA) tasks may be susceptible to text-only or superficial visual shortcuts, leading to unreliable assessments of visual understanding. To address these limitations, we introduce a benchmark and training framework for shortcut-resistant cross-scale pathology reasoning. We design an Adversarial Text-only Screening strategy for semantic reasoning questions and a Structure-controlled Distractor Sampling strategy for visual grounding questions, encouraging models to rely on cross-scale visual evidence. Based on this pipeline, we construct PathScale-VQA, a high-quality cross-scale pathology VQA benchmark with 10,373 multiple-choice questions grounded in 1,368 diagnostic paths across multiple magnification levels. Building on the semantic reasoning set, PathScale-R1 is optimized through Difficulty-driven Reasoning Distillation supervised fine-tuning followed by reinforcement learning with a Scale-aware Reasoning Structure reward, which encourages the use of evidence across magnifications. Extensive experiments demonstrate state-of-the-art performance of PathScale-R1 on cross-scale reasoning tasks and effective transfer to conventional single-scale pathology VQA. Our code is available at https://github.com/iMVR-PL/PathScale-R1.
Chi Phan, Tianyi Zhang, Yufeng Wu +7
Department of Electrical and Computer Engineering, National University of Singapore, Singapore 117417 · Department of Biomedical Engineering, National University of Singapore, Singapore 117417, Singapore · Department of Pathology, Fujian Medical University Cancer Hospital & Fujian Cancer Hospital, Fuzhou, China +1