What Makes High-Magnification Knowledge Transferable? A Study of Cross-Resolution Distillation in Whole-Slide Imaging
Organizations: Department of Computer Science and Software Engineering (CSSE) Concordia University Montreal, Canada · Department of Computer Science and Software Engineering (CSSE) Concordia University Montreal, Canada Mila – Quebec AI Institute Montreal, Canada
Abstract
Cross-resolution knowledge distillation aims to improve low-magnification whole- slide analysis by transferring high-magnification representations, yet the conditions for useful transfer remain unclear. We develop a decomposition-based analysis of teacher access, representation loss, and model excess, motivating three questions: whether (a) teacher targets help the task, (b) low-magnification students can predict them, and (c) slide models benefit from those predictions. We investigate them through controlled experiments across ten pathology cohorts spanning classifi- cation, grading, and survival prediction. In the main comparison, providing teacher regional means alongside native low-magnification features improves downstream performance in all ten cohorts. Direct prediction achieves lower reconstruction error than residual prediction, yet the predicted features underrepresent variation in the teacher targets. Moreover, better reconstruction does not consistently improve downstream scores, and retaining native features changes performance even when the predicted teacher features are held fixed. Together, these findings expose a gap between reconstructing teacher representations and realizing their downstream value. They challenge the sufficiency of reconstruction error as a measure of cross-resolution transfer and provide a diagnostic framework for examining where that transfer breaks down. Future distillation designs must account for both what students can predict and how slide models use those predictions.
Figures & tables
Appendix figures & tables34 assets
Supplementary material from the paper’s appendix.
Appendix
| Dataset | Metric | Slides/images | Group | Train / val. / test |
|---|---|---|---|---|
| CAM16 | AUC | 270 | Slide | 194 / 22 / 54 |
| CAM17 | Macro | 499 | Challenge group | 70 / 10 / 20 |
| Private-CRC | Macro | 837 | Slide | 604 / 67 / 166 |
| BRACS | Macro | 546 | Patient | 135 / 17 / 37 |
| PANDA | Quadratic | 10,615 | Image | 7645 / 847 / 2123 |
| TCGA-KIRC | Case C-index | 494 | Case | 337 / 52 / 101 |
| Dataset | Epochs | Learning rate | Decay | Accum. | Selection |
|---|---|---|---|---|---|
| CAM16 | 15 | 16 | Macro | ||
| CAM17 | 15 | 16 | Macro | ||
| Private-CRC | 15 | 16 | Macro | ||
| BRACS | 15 | 16 | Macro | ||
| PANDA | 20 | 16 | Quadratic | ||
| KIRC | 20 | 1 | Case C-index |
| input | teacher construction | |||
|---|---|---|---|---|
| Dataset / metric | Native | Regional mean | Repeated mean | Individual |
| Tokens per region | 1 | 1 | 16 | 16 |
| CAM16 / AUC | ||||
| CAM17 / | ||||
| Private-CRC / | ||||
| PANDA / | ||||
| Dataset | Native | Native + mean | Native + mean and spatial | Native + mean and direction control |
|---|---|---|---|---|
| CAM16 | ||||
| CAM17 | ||||
| Private-CRC | ||||
| PANDA | ||||
| BRACS | ||||
| KIRC |
| Dataset | Native | Native + mean | Native + mean and spatial | Native + mean and direction control |
|---|---|---|---|---|
| CAM16 | ||||
| CAM17 | ||||
| Private-CRC | ||||
| PANDA | ||||
| BRACS | ||||
| KIRC |
| Primary metric, higher is better | ||||||
|---|---|---|---|---|---|---|
| Dataset | Native | Teacher mean | Teacher spatial | Predicted mean | Predicted spatial | Initial spatial |
| CAM16 | ||||||
| CAM17 | ||||||
| Private-CRC | ||||||
| PANDA | ||||||
| BRACS | ||||||
| Regional mean | Nonconstant components | |||
|---|---|---|---|---|
| Excluded dataset | Predictor | Constant | Predictor | Constant |
| CAM16 | ||||
| CAM17 | ||||
| Private-CRC | ||||
| PANDA | ||||
| BRACS | ||||
| Dataset | Metric | Repeated mean | Spatial summary | Permuted summary |
|---|---|---|---|---|
| CAM16 | AUC | |||
| CAM17 | ||||
| Private-CRC | ||||
| PANDA | ||||
| BRACS | ||||
| KIRC | C-index |
| TransMIL | Gated ABMIL | DSMIL | RRT-MIL | |||||
|---|---|---|---|---|---|---|---|---|
| Dataset | Native | Mean | Native | Mean | Native | Mean | Native | Mean |
| CAM16 | ||||||||
| CAM17 | ||||||||
| Private-CRC | ||||||||
| PANDA | ||||||||
| BRACS | ||||||||
| Native | Constant | Direct | Residual | |||
|---|---|---|---|---|---|---|
| Dataset | Zero | Random | Zero | Random | ||
| CAM16 | ||||||
| CAM17 | ||||||
| Private-CRC | ||||||
| PANDA | ||||||
| BRACS | ||||||
| Zero final-layer initialization | ||||
|---|---|---|---|---|
| Direct | Residual | |||
| Dataset | Permuted | Aligned | Permuted | Aligned |
| CAM16 | ||||
| CAM17 | ||||
| Private-CRC | ||||
| PANDA | ||||
| MSE | |||||
|---|---|---|---|---|---|
| Native | Direct | Residual | |||
| Dataset | Initial | Trained | Initial | Trained | |
| CAM16 | |||||
| CAM17 | |||||
| Private-CRC | |||||
| PANDA | |||||
| Centered variance ratio | Pearson correlation of pairwise distances | ||||
|---|---|---|---|---|---|
| Dataset | Direct | Residual | Native | Direct | Residual |
| CAM16 | |||||
| CAM17 | |||||
| Private-CRC | |||||
| PANDA | |||||
| BRACS | |||||
| Dataset | Regions | Native | Direct | Residual |
|---|---|---|---|---|
| CAM16 | 128 | |||
| CAM17 | 135 | |||
| Private-CRC | 149 | |||
| PANDA | 52 | |||
| BRACS | 134 | |||
| KIRC | 128 |
| A. Independent evaluation within distillation datasets | ||||
|---|---|---|---|---|
| Regional mean | Spatial components | |||
| Excluded dataset | Constant | Network | Constant | Network |
| CAM16 | ||||
| CAM17 | ||||
| Private-CRC | ||||
| PANDA | ||||
| Independent evaluation within distillation datasets | ||||
|---|---|---|---|---|
| Regional mean | Spatial components | |||
| Excluded dataset | Base | Corrected | Base | Corrected |
| CAM16 | ||||
| CAM17 | ||||
| Private-CRC | ||||
| PANDA | ||||
| Dataset | Metric | Global input | Spatial input | ||
|---|---|---|---|---|---|
| CAM16 | AUC | ||||
| CAM17 | |||||
| Private-CRC | |||||
| PANDA | |||||
| BRACS | |||||
| KIRC | C-index |
| Dataset / metric | Native Zero augmentation | Native Initial prediction | Native Trained mean + spatial | Native Same trained mean |
|---|---|---|---|---|
| CAM16 / AUC | ||||
| CAM17 / | ||||
| Private-CRC / | ||||
| PANDA / | ||||
| BRACS / | ||||
| KIRC / C-index |
| Dataset / metric | Native | Direct Zero init. | Residual Zero init. | Direct Random init. | Residual Random init. |
|---|---|---|---|---|---|
| CAM16 / AUC | |||||
| CAM17 / | |||||
| Private-CRC / | |||||
| PANDA / | |||||
| BRACS / | |||||
| KIRC / C-index |
| Dataset / metric | Native | Direct | Direct + native | Residual | Residual + native |
|---|---|---|---|---|---|
| Zero final-layer initialization before distillation | |||||
| CAM16 / AUC | |||||
| CAM17 / | |||||
| Private-CRC / | |||||
| PANDA / | |||||
| BRACS / | |||||
| Dataset | Native | Direct | Direct + native | Residual | Residual + native |
|---|---|---|---|---|---|
| Zero final-layer initialization before distillation | |||||
| CAM16 | |||||
| CAM17 | |||||
| Private-CRC | |||||
| PANDA | |||||
| BRACS | |||||
| Dataset / metric | Frozen UNI2-h | Mean KD Encoder | Mean KD Mean head | Mean KD + flow Regional mean | Mean KD + flow Individual feature |
|---|---|---|---|---|---|
| CAM16 / AUC | |||||
| CAM17 / | |||||
| Private-CRC / | |||||
| PANDA / | |||||
| BRACS / | |||||
| KIRC / C-index |
| Dataset / metric | Flow toward regional mean | Flow toward individual feature |
|---|---|---|
| CAM16 / AUC | ||
| CAM17 / | ||
| Private-CRC / | ||
| PANDA / | ||
| BRACS / | ||
| KIRC / C-index |