Foundation models trained on large pathology image corpora now provide strong, transferable representations for computational pathology. Over the past few years a series of such models has been released, each trained on more slides than the last; on standard classification and segmentation benchmarks, the leading models are now separated by small margins. In clinical use, however, the foundation model is applied to images from hospitals, scanners, and staining protocols outside its training data. Encoders generally embed these acquisition factors alongside biological information, which may introduce downstream errors and hinder safe clinical adoption. A pathology foundation model should therefore be robust to acquisition shift without giving up representation quality, yet robustness is seldom the axis along which models are compared. In this report, we introduce HERO (Histology Encoder for Robust Representation in Oncology), a ViT-G/14 pathology foundation model trained with the DINO and iBOT objectives and refined with high-resolution Gram anchoring on a morphology-balanced corpus of 500 million tiles from approximately 575,000 clinical whole-slide images. Across the evaluated public benchmarks, HERO shows the strongest robustness to center, scanner, and stain variation among the compared state-of-the-art foundation models, performs comparably on tile-level classification, segmentation, and gene-expression prediction, ranks first on average across 39 evaluated slide-level clinical tasks, and, under an equal-weighted framework-level analysis, has the best average rank across the six benchmark frameworks.
Figures & tables
Figure 1: Robustness and overall standing versus pretraining dataset size. (a) PathoROB Robustness Index (higher is better). The gray dashed line is a log-linear fit to the six baselines. (b) Average rank across the six benchmark frameworks, each weighted equally (lower is better). Colored dashed lines mark the strongest public baseline. Per-framework ranks are in Table 13 .
Figure 2: Composition of the HERO pretraining archive (approximately 575,000 WSIs; 1.126 billion retained 20 × tiles). (a) Share of slides and of retained tiles by specimen organ, with counts annotated. (b) The same by disease lineage; the 38 smallest lineages are pooled as Other . (c) Scanner composition by slide and by tile.
Figure 3: HERO workflow. (a) Curation: tiles are embedded with a frozen DINOv2 encoder, clustered at two levels, and sampled under coarse-cluster quotas. (b) Representation extraction: normalized CLS token, mean patch token, and their concatenation (CLS+Mean). (c) Training: Stage 1 uses DINO and iBOT objectives; Stage 2 adds Gram anchoring.
Model
# WSIs
Data Access
Camelyon Breast
TCGA 2 × 2 Multi-organ
Tolkach ESCA Esophagus
Average RI ↑
Virchow2
3.1M
Private
0.806
0.822
0.955
0.861
H-Optimus-1
1M
Private
0.645
0.853
0.944
0.814
UNI2
350K
Private
0.544
0.803
0.923
0.757
Prov-GigaPath
171K
Private
0.399
0.738
0.754
0.630
UNI
100K
Private
0.145
0.747
0.902
0.598
Phikon-v2
60K
Public
0.019
0.619
0.768
0.469
Table 1: PathoROB results. Robustness Index (higher means neighbor structure follows biology rather than center). #WSIs: reported pretraining slide count.
Model
Embedding consistency Cosine similarity ↑
Scanner Top-10 accuracy ↑
Stain Top-10 accuracy ↑
Combined Top-10 accuracy ↑
Average Metric ↑
H-Optimus-0
0.685
0.744
0.327
0.166
0.480
Virchow2
0.777
0.609
0.306
0.163
0.464
Prov-GigaPath
0.570
0.592
0.118
0.054
0.333
UNI2-h
0.591
0.501
0.190
0.046
0.332
UNI
0.547
0.532
0.169
0.053
0.325
Phikon-v2
0.557
0.064
0.030
0.003
0.164
Table 2: PLISM results. Median over matched slide pairs for each acquisition axis.
Model
Tile-level
Slide-level
Average
PCam 10 Bal. acc. ↑
BACH Bal. acc. ↑
BRACS Bal. acc. ↑
CRC Bal. acc. ↑
PCam Bal. acc. ↑
Gleason Bal. acc. ↑
CoNSeP Dice ↑
MoNuSAC Dice ↑
Cam16 Bal. acc. ↑
PANDA Bal. acc. ↑
UNI2
0.887
0.915
0.661
0.965
0.950
0.775
0.630
0.642
0.849
0.657
0.793
Virchow2
0.851
0.883
0.624
0.967
0.938
0.783
0.640
0.669
0.861
0.646
0.786
H-Optimus-0
0.824
0.759
0.615
0.955
0.943
0.770
0.644
0.685
0.827
0.671
0.769
Prov-GigaPath
0.852
0.759
0.616
0.951
0.945
0.724
0.626
0.680
0.815
0.653
0.762
UNI
0.815
0.785
0.593
0.944
0.937
0.750
0.628
0.659
0.833
0.659
0.760
Table 3: EVA results. Balanced accuracy for classification and Dice for segmentation (CoNSeP, MoNuSAC) over eight tile-level and two slide-level tasks.
Model
kNN F1 ↑
Linear probe F1 ↑
Few-shot F1 ↑
Segmentation Dice ↑
Calibration ECE (%) ↓
Adversarial F1 drop (%) ↓
UNI2
0.833
0.857
0.798
0.690
3.9
31.7
Virchow2
0.829
0.848
0.739
0.693
3.9
31.1
UNI
0.808
0.835
0.781
0.678
3.8
40.3
H-Optimus-0
0.814
0.838
0.762
0.652
4.0
43.9
H-Optimus-1
0.825
0.851
0.773
0.645
3.5
57.4
Prov-GigaPath
0.795
0.829
0.755
0.635
3.4
42.1
Table 4: THUNDER results. Frozen-feature prediction, calibration, and adversarial stress. Adversarial F1 drop is the clean-to-adversarial decrease at the benchmark budget. Parentheses: HERO’s within-table rank.
Model
IDC Breast
PRAD Prostate
PAAD Pancreas
SKCM Skin
COAD Colon
READ Rectum
ccRCC Kidney
LUAD Lung
LYMPH-IDC Lymph node
Average Pearson r↑
H-Optimus-1
0.602
0.378
0.496
0.659
0.320
0.242
0.253
0.578
0.277
0.423
H-Optimus-0
0.598
0.385
0.491
0.645
0.309
0.222
0.268
0.559
0.259
0.415
UNI2
0.590
0.357
0.500
0.661
0.301
0.222
0.264
0.559
0.273
0.414
Virchow2
0.597
0.353
0.478
0.640
0.258
0.207
0.272
0.569
0.257
0.403
Prov-GigaPath
0.551
0.370
0.475
0.562
0.299
0.196
0.243
0.541
0.250
0.387
UNI
0.589
0.294
0.481
0.635
0.261
0.184
0.240
0.546
0.256
0.387
Table 5: HEST results. Pearson correlation for the 50 most variable genes in each task.
Figure 4: Slide-level comparison on Patho-Bench (39 tasks). (a) Category and overall averages for HERO and six public encoders. (b) Per-task scores, color-scaled within each task (blue: low; red: high); best per task in bold. The last two columns give each model’s average over the 39 tasks and the number of tasks on which it scored best.
Appendix figures & tables8 assets
Supplementary material from the paper’s appendix.
Appendix
Shared across both stages
Backbone
ViT-G/14, 4 register tokens, 1,536-d, SwiGLU
Parameters
≈ 1.1B
Precision
bfloat16, FSDP
Drop-path rate
0.4
DINO / iBOT prototypes
131,072 / 131,072
KDE weight
0.05
Appendix
Table 6: HERO training configuration. Upper block: settings shared by both stages. Lower block: stage-specific settings; a dash marks settings not used in Stage 1.
Framework
Level
Datasets
Property measured
Protocol
Metric
PathoROB
Tile
3
Center robustness
kNN neighborhood
Robustness Index
PLISM
Tile
1 a
Scanner and stain robustness
Retrieval
Cosine sim., Top-10 acc.
EVA
Tile + slide
10 b
Classification, segmentation
Linear probe, ABMIL
Bal. acc., Dice
THUNDER
Tile
20
Classification, segmentation, stress
kNN, linear probe, few-shot
F1, Dice, ECE, F1 drop
HEST
Tile
9
Gene-expression prediction
Ridge regression
Pearson r
Patho-Bench
Slide
19 c
Clinical slide-level tasks
ABMIL
AUROC, Bal. acc., C-index
Appendix
Table 7: Evaluation benchmarks. Datasets counts the constituent datasets or cohorts. Higher is better for all metrics except the two THUNDER stress metrics.
Category
Cohort
Tasks
Morphological subtyping (4)
BRACS
slide-level coarse; slide-level fine
EBRAINS
slide-level coarse; slide-level fine
Tumor grading (2)
IMP-CRS
grade
PANDA
ISUP grade
Mutation prediction (21)
CPTAC-BRCA
PIK3CA; TP53
CPTAC-CCRCC
BAP1; PBRM1
Appendix
Table 8: The 39 Patho-Bench tasks used in Fig. 4 . Cohort and task names follow the Patho-Bench naming.
Training recipe
EVA Average ↑
HEST Pearson r↑
PathoROB RI ↑
Stage 1 only (DINO/iBOT)
0.776
0.414
0.881
+ Gram anchoring
0.780
0.404
0.892
Δ (%)
↑ 0.5%
↓ 2.4%
↑ 1.2%
Appendix
Table 9: Effect of Stage 2 Gram anchoring on benchmark averages. Stage 1 only is the DINO/iBOT checkpoint at 450k steps; + Gram anchoring continues it for 20k steps with the Gram-anchoring loss and is the released HERO model. Δ : relative change from Stage 1 only to + Gram anchoring, ↑ green for an increase and ↓ red for a decrease. The released configuration is shaded.
Training recipe
Camelyon Breast
TCGA 2 × 2 Multi-organ
Tolkach ESCA Esophagus
Average RI ↑
Stage 1 only
0.813
0.873
0.956
0.881
+ Gram anchoring
0.836
0.884
0.956
0.892
Δ (%)
↑ 2.8%
↑ 1.3%
0.0%
↑ 1.2%
Appendix
Table 10: PathoROB with and without Gram anchoring. Robustness Index as in Table 1 . Δ : relative change from Stage 1 only to + Gram anchoring, ↑ green for an increase and ↓ red for a decrease. The released configuration is shaded.
Training recipe
Tile-level
Slide-level
Average
PCam 10
BACH
BRACS
CRC
PCam
Gleason
CoNSeP
MoNuSAC
Cam16
PANDA
Stage 1 only
0.866
0.832
0.629
0.967
0.945
0.776
0.633
0.644
0.838
0.633
0.776
+ Gram anchoring
0.875
0.848
0.622
0.963
0.943
0.793
0.633
0.657
0.834
0.636
0.780
Δ (%)
↑ 1.0%
↑ 1.9%
↓ 1.1%
↓ 0.4%
↓ 0.2%
↑ 2.2%
0.0%
↑ 2.0%
↓ 0.5%
↑ 0.5%
↑ 0.5%
Appendix
Table 11: EVA with and without Gram anchoring. Metrics as in Table 3 . Δ : relative change from Stage 1 only to + Gram anchoring, ↑ green for an increase and ↓ red for a decrease. The released configuration is shaded.
Training recipe
IDC Breast
PRAD Prostate
PAAD Pancreas
SKCM Skin
COAD Colon
READ Rectum
ccRCC Kidney
LUAD Lung
LYMPH-IDC Lymph node
Average Pearson r↑
Stage 1 only
0.572
0.374
0.477
0.666
0.303
0.216
0.293
0.562
0.261
0.414
+ Gram anchoring
0.585
0.378
0.489
0.614
0.261
0.209
0.271
0.561
0.269
0.404
Δ (%)
↑ 2.3%
↑ 1.1%
↑ 2.5%
↓ 7.8%
↓ 13.9%
↓ 3.2%
↓ 7.5%
↓ 0.2%
↑ 3.1%
↓ 2.4%
Appendix
Table 12: HEST with and without Gram anchoring. Metric as in Table 5 . Δ : relative change from Stage 1 only to + Gram anchoring, ↑ green for an increase and ↓ red for a decrease. The released configuration is shaded.
Model
Average rank within each benchmark
Average rank across frameworks
PathoROB
PLISM
EVA
THUNDER
HEST
Patho-Bench
HERO
1.00
1.00
3.35
4.00
3.44
2.86
2.61
H-Optimus-0/1
2.67
2.25
3.50
4.08
1.61
4.10
3.04
Virchow2
2.33
2.75
2.70
3.08
3.67
4.56
3.18
UNI2
4.00
5.25
2.40
2.00
2.44
3.18
3.21
UNI
5.33
5.50
4.90
4.00
5.39
3.72
4.81
Appendix
Table 13: Average rank across benchmark frameworks. Models are ranked (1 = best; ties share the average rank) on every metric a framework reports, that is, each column of Tables 1 – 5 and each of the 39 Patho-Bench tasks. Ranks are averaged within each framework, and the last column averages the six frameworks with equal weight. Metrics per framework: PathoROB 3, PLISM 4, EVA 10, THUNDER 6, HEST 9, Patho-Bench 39. H-Optimus-0 and H-Optimus-1 are pooled as one entry. Best in bold, second best underlined.
Pathology foundation models (FMs) produce powerful tile-level representations which remain sensitive to scanner and staining variability, undermining deployment across laboratories. We develop a novel fine-tuning recipe that improves the robustness of pathology FMs to acquisition factors. Applied to ten different FMs, our fine-tuning strategy consistently improves robustness for every model as well as downstream performance, with no observed trade-off. On average, it raises the PathoROB robustness index by 23% (from 0.72 to 0.87) and increases the overall cross-benchmark performance by 43% on Patho-Bench, HEST and THUNDER combined, with individual gains reaching up to 72% in robustness (Phikon-v2) and 76% in performance (Midnight-12k). We publicly release the fine-tuned versions of Phikon-v2 (Phaet) and Midnight-12k (Mascaret) at https://huggingface.co/wearewaiv/models.
Alexandre Filiot, Oskar Thaeter, Benoit Schmauch +1
1Waiv · Institute of Pathology, Technical University of Munich · School of Computation, Information and Technology, Technical University of Munich
Foundation Models (FMs) have recently redefined the state-of-the-art in histopathology by providing robust representations for whole-slide image (WSI) analysis. However, selecting the optimal foundation model (FM) for a specific clinical cohort currently requires multiple preprocessing steps, followed by computationally expensive feature extraction and the training of a Multiple Instance Learning (MIL) aggregator for every model. In this work, we investigate whether efficient tile-level linear probing can serve as a reliable proxy for slide-level performance, reducing the need to run full slide-level pipelines for every candidate encoder. We benchmark 19 state-of-the-art FMs on 42 slide-level and 16 tile-level tasks, comparing tile probing metrics against slide-level outcomes using ABMIL and Mean Pooling aggregations. We observe a high correlation between tile and slide performance across varying task difficulties, indicating that encoder representation quality is the primary determinant of WSI success. Sensitivity analyses show that transferability is stable across models and is more influenced by cohort sizes and numbers of tiles per slide than by average task difficulty. We also measure the agreement in best performing models between tile and slide-level tasks, showing tile benchmarks reliably shortlist strong candidates. Overall, our study indicates that tile-level benchmarking provides an efficient and practical first step for narrowing down candidate models, while slide-level evaluation remains essential for final validation on clinical tasks.
Sofiène Boutaj, Leo Fillioux, Maria Vakalopoulou +2
Université Paris-Saclay, CentraleSupélec, Gustave Roussy, INSERM, IHU PRISM, Cancer Data Science Unit, France · Université Paris-Saclay, CentraleSupélec, MICS Laboratory, France
Foundation models have emerged as a driving force in computational pathology, with the potential to transform cancer diagnosis, prognosis, and treatment selection by learning transferable representations from large-scale histopathology data. A growing landscape of pathology foundation models now spans diverse data sources, architectures, and downstream applications. However, most pretrained models operate only at the image-tile level, use restrictive licenses, and remain computationally expensive, limiting large-scale slide-level clinical and research use. Here, we introduce GigaPath-Flash and GigaTIME-Flash, efficient models for whole-slide pathology AI and spatial proteomics prediction. GigaPath-Flash combines a 22M-parameter ViT-S tile encoder with a 21M-parameter LongNet slide encoder, both pretrained on large-scale real-world histopathology data. Its compact tile encoder is distilled from the billion-parameter GigaPath (ViT-g) teacher and shared by both models. GigaPath-Flash retains 97% of GigaPath's average slide-level performance with 50x less compute. GigaTIME-Flash extends this backbone to predict the tumor immune microenvironment directly from routine H&E images. It surpasses the original CNN-based GigaTIME in prediction quality while running 6x faster and using 8x less GPU memory. Together with GigaPath and GigaTIME, these models form an open-weight, Apache-2.0-licensed family pretrained on large-scale real-world clinical data. By releasing all models and weights, we provide accessible building blocks for computational pathology, immuno-oncology, and precision health.
Naoto Usuyama, Jeya Maria Jose Valanarasu, Sicong Yao +27
Microsoft Research · Paul G. Allen School of Computer Science and Engineering, University of Washington · Providence Genomics +2