See, Measure, and Reason: Learning Visually Grounded Reasoning in Pathology
Organizations: College of Computer Science, Sichuan University · Department of Pathology and Institute of Clinical Pathology, West China Hospital, Sichuan University · Department of Computer Science, School of Computing, National University of Singapore · School of Intelligent Systems Engineering, Sun Yat-sen University
Abstract
Pathological assessment relies on recognizing fine-grained visual details in histological images. Vision-language models (VLMs) increasingly support pathology interpretation, yet their ability to perceive these details remains inadequate. This weakness leads to inaccurate cellular observations that can persist even when final answers are correct. In this paper, we propose ASPECT to improve visually grounded reasoning through explicit supervision of cellular appearance and abundance. ASPECT trains intermediate visual tokens through pathology feature reconstruction, cell feature alignment, and count supervision. Three-stage supervised fine-tuning teaches the model to perceive, generate visual tokens, and reason, followed by reinforcement learning that rewards answer correctness and consistency with reported measurements. We also introduce PathoVernier, a benchmark of 759 expert-reviewed questions from five pathology datasets covering four cellular composition tasks. It evaluates both final answers and intermediate measurements to expose errors hidden by answer accuracy. On PathoVernier, ASPECT achieves relative accuracy gains of approximately 19.2% over the strongest baseline, Gemini-3.1-Pro, and 99.3% over its Qwen3-VL-8B backbone, while reducing RAWR, which measures counting errors within correct responses, by 28.1% and 42.7%, respectively. ASPECT also improves over its backbone on three external pathology benchmarks covering classification and question answering beyond cellular composition tasks.
Figures & tables
| Models | PathoVernier | PathCLS | PathVQA | Quilt-VQA | ||
|---|---|---|---|---|---|---|
| Acc | CA | RAWR | Acc | Acc | Acc | |
| Closed-source Models | ||||||
| GPT-5.5 ( OpenAI, 2026a ) | 0.59 | 0.65 | 0.76 | 0.55 | 0.69 | 0.69 |
| Gemini-3.1-Pro ( Google DeepMind, 2026 ) | 0.63 | 0.71 | 0.68 | 0.67 | 0.78 | 0.65 |
| Open-source Models | ||||||
| Qwen3-VL-8B ( Bai et al., 2025 ) | 0.37 | 0.53 | 0.85 | 0.37 | 0.67 | 0.58 |
| Configuration | Acc | CA | RAWR | Count Acc |
|---|---|---|---|---|
| ASPECT-SFT | 0.70 | 0.80 | 0.53 | 0.44 |
| w/o visual supervision | 0.64 | 0.77 | 0.59 | 0.32 |
| w/o pathology feature reconstruction | 0.68 | 0.75 | 0.56 | 0.27 |
| w/o cell feature alignment | 0.67 | 0.77 | 0.54 | 0.29 |
| w/o count supervision | 0.72 | 0.78 | 0.57 | 0.36 |
| Direct Reason training (Stage 3) | 0.62 | 0.74 | 0.56 | 0.31 |
| Configuration | RL Reward | Acc | CA | RAWR | Count Acc | Consistency |
|---|---|---|---|---|---|---|
| ASPECT-SFT | N/A | 0.70 | 0.80 | 0.53 | 0.44 | 0.94 |
| Answer only | Acc | 0.72 | 0.79 | 0.55 | 0.43 | 0.92 |
| Consistency only | 0.5 Con | 0.69 | 0.76 | 0.51 | 0.41 | 0.98 |
| Answer + count score | Acc + 0.5 S | 0.74 | 0.80 | 0.50 | 0.45 | 0.96 |
| ASPECT-8B | Acc + 0.5 Con | 0.75 | 0.81 | 0.49 | 0.47 | 0.99 |
Appendix figures & tables18 assets
Supplementary material from the paper’s appendix.
Appendix
| Source | SFT | RL | Validation | PathoVernier | |
| Visual | VQA | ||||
| Lizard ( Graham et al., 2021 ) | 2,520 | 2,944 | 559 | 40 | 400 |
| PUMA ( Schuiveling et al., 2025 ) | 2,745 | 1,010 | 219 | 5 | 142 |
| PanNuke ( Gamper et al., 2020 ) | 6,780 | 2,565 | 567 | 20 | 139 |
| CoNSeP ( Graham et al., 2019 ) | 196 | 133 | 27 | 2 | 44 |
| NuCLS ( Amgad et al., 2022 ) | 854 | 77 | 0 | 26 | 34 |
| Source | Original label(s) | Target category/categories |
| Lizard | Epithelial | Tumor; normal epithelial ∗ |
| Connective tissue | Stromal-like | |
| Lymphocyte | Lymphocyte | |
| Plasma, neutrophil, eosinophil | Other inflammatory | |
| PUMA | Tumor | Tumor |
| Epithelium | Normal epithelial |
| Source | Label source | Feature-alignment targets | Count supervision |
|---|---|---|---|
| NCT-CRC-HE-100K | Patch-level tissue labels | Mapped patch category for all detected nuclei | Detected nucleus count assigned to that category |
| PathVQA | Native CellViT PanNuke predictions | Mapped nucleus categories; shared inflammatory targets | Predicted category counts; aggregate inflammatory count |
| Skill | Questions | Skill | Questions |
| Multi-step composition | 388 | Count comparison | 85 |
| Region selection | 407 | Region comparison | 60 |
| Dominant cell type | 208 | Cell-type comparison | 60 |
| Count band prediction | 124 | Yes/no | 40 |
| Total | 1,372 | ||
| Source | Region selection | Region comparison | Cell-type comparison | Multi-step composition | Total |
|---|---|---|---|---|---|
| Lizard | 82 | 85 | 135 | 98 | 400 |
| PUMA | 45 | 47 | 10 | 40 | 142 |
| PanNuke | 35 | 36 | 33 | 35 | 139 |
| CoNSeP | 17 | 7 | 13 | 7 | 44 |
| NuCLS | 11 | 11 | 2 | 10 | 34 |
| Total | 190 | 186 | 193 | 190 | 759 |
| Organ | Questions | Organ | Questions |
|---|---|---|---|
| Colon | 480 | Kidney | 4 |
| Skin | 142 | Ovary | 3 |
| Breast | 60 | Liver | 2 |
| Bile duct | 15 | Pancreas | 2 |
| Head and neck | 11 | Prostate | 2 |
| Uterus | 10 | Stomach | 2 |
| Source | Resolution ( m/pixel) | Patch width ( m) | Approx. magnification | Questions |
| Lizard | 0.50 | 128.0 | 400 | |
| PanNuke, CoNSeP | 0.25 | 64.0 | 183 | |
| PUMA | 0.22 | 56.3 | 142 | |
| NuCLS | 0.20 | 51.2 | 34 | |
| Total | 759 | |||
| Setting | Value |
|---|---|
| Model and adaptation | |
| Backbone | Qwen3-VL-8B |
| Pathology feature tokens | 8 |
| Cell tokens | 6 |
| Visual teachers | UNI; CellViT-SAM-H |
| Teacher updates | Frozen; targets extracted offline |
| Benchmark | Evaluation subset | Questions | Local token limit | API token limit |
|---|---|---|---|---|
| PathoVernier | Full benchmark | 759 | 512 | 512 |
| PathCLS | Official test | 1,632 | 512 | 512 |
| PathVQA | H&E, closed-ended test | 1,306 | 512 | 512 |
| Quilt-VQA | H&E, closed-ended test | 314 | 512 | 512 |
| Partition | Stromal-like | Lymphocyte | Epithelial-like | Inflammatory |
|---|---|---|---|---|
| Quadrants | ||||
| Horizontal bands | ||||
| Vertical bands |
| Model | CA coverage | Complete-count coverage | Count Acc | |
| GPT-5.5 | 748/759 | 749/759 | 447 | 0.210 |
| Gemini-3.1-Pro | 733/759 | 750/759 | 472 | 0.296 |
| Qwen3-VL-8B | 233/759 | 261/759 | 106 | 0.068 |
| Qwen3-VL-32B | 682/759 | 700/759 | 333 | 0.178 |
| InternVL3-8B | 448/759 | 449/759 | 166 | 0.071 |
| InternVL3-38B | 750/759 | 753/759 | 340 | 0.153 |
| Model | Acc | CA | RAWR |
|---|---|---|---|
| Qwen3-VL-8B | 0.374 [0.341, 0.408] | 0.532 [0.471, 0.595] | 0.854 [0.801, 0.901] |
| Gemini-3.1-Pro | 0.626 [0.588, 0.663] | 0.715 [0.681, 0.749] | 0.680 [0.651, 0.708] |
| GPT-5.5 | 0.594 [0.558, 0.631] | 0.652 [0.616, 0.689] | 0.756 [0.728, 0.784] |
| ASPECT-8B | 0.746 [0.714, 0.776] | 0.807 [0.777, 0.836] | 0.489 [0.460, 0.518] |
| Comparison | Acc gain | CA gain | RAWR reduction |
| ASPECT vs Qwen3-VL-8B | +0.372 [+0.327, +0.415] | +0.275 [+0.207, +0.342] | +0.365 [+0.305, +0.420] |