When Less Data Favors Smaller Teachers: Rethinking Teacher Capacity and Data Selection for Knowledge Distillation
Organizations: Pohang University of Science and Technology (POSTECH)
Abstract
Data pruning reduces the training cost of knowledge distillation (KD). However, the preferred teacher capacity changes with the data budget: smaller teachers can outperform larger ones when limited training data are available. Understanding what drives this shift is important not only for teacher choice but also for identifying which samples are useful for distillation. We analyze teacher supervision by decomposing it into relational ordering---the ranking of classes---and score geometry---the magnitudes and margins of class probabilities---and show that the small-teacher advantage in the low-data regime arises not only from score geometry but also from relational ordering. Beyond understanding teacher capacity, our analysis reveals two properties of effective subsets: samples should match the difficulty appropriate for the available budget, and their relational signals should be diverse rather than redundant. Based on these findings, we propose DVA (Difficulty- and Volume-Aware data selection for KD), a training-dynamics-free method, which uses a small teacher as a proxy for budget-aware difficulty filtering and class-conditional relational volume maximization. Despite requiring no training dynamics statistics, our method remains competitive with training-dynamics-based methods while consistently outperforming training-dynamics-free baselines.
Figures & tables
| Relational Ordering | Score Geometry | Drop Rate | |
| 0.1 | 0.9 | ||
| Large | Large | 78.82 | 46.93 |
| Large | Small | 78.62 (-0.20) | 55.11 (+8.18) |
| Small | Large | 78.97 (+0.15) | 50.75 (+3.82) |
| Small | Small | 75.33 (-3.49) | 62.49 (+15.56) |
| Method | Drop Rate | |||||
| 0.1 | 0.3 | 0.5 | 0.7 | 0.8 | 0.9 | |
| Training-Dynamics-Based | ||||||
| Forgetting | 79.35 0.27 | 77.69 0.29 | 73.56 0.37 | 61.45 0.22 | 51.02 0.51 | 35.06 0.73 |
| EL2N | 79.06 0.28 | 77.21 0.25 | 70.23 0.30 | 50.85 0.68 | 35.02 0.60 | 16.44 0.29 |
| CCS | 79.13 0.23 | 7 7.81 0.23 | 74.81 0.21 | 69.82 0.36 | 64.84 0.35 | 53.77 0.62 |
| D 2 | 79.02 0.27 | 77.58 0.20 | 7 4.96 0.17 | 6 9.99 0.27 | 66.14 0.34 | 5 4.49 0.63 |
| Drop | Random | Filtered Rand. | DVA (Ours) | Gain |
| 0.7 | 67.07 0.49 | 68.09 0.20 | 69.09 0.23 | +1.00 |
| 0.8 | 61.35 0.43 | 63.40 0.32 | 64.45 0.36 | +1.05 |
| 0.9 | 46.93 0.80 | 53.02 0.62 | 53.76 0.77 | +0.74 |
| Drop Rate | ||||||
| Proxy Teacher | 0.1 | 0.3 | 0.5 | 0.7 | 0.8 | 0.9 |
| Width-4 | 79.14 0.19 | 77.06 0.28 | 73.84 0.26 | 67.85 0.18 | 63.17 0.30 | 53.31 0.67 |
| Width-8 (small) | 79.13 0.17 | 77.43 0.20 | 74.62 0.16 | 69.09 0.23 | 64.45 0.36 | 53.76 0.77 |
| Width-16 | 79.09 0.22 | 76.99 0.18 | 73.77 0.27 | 67.91 0.20 | 63.05 0.32 | 50.30 1.19 |
| Width-32 | 79.00 0.22 | 76.74 0.33 | 73.42 0.26 | 65.94 0.29 | 59.48 0.72 | 42.05 1.38 |
| Width-64 (large) | 79.12 0.28 | 76.62 0.20 | 73.25 0.45 | 64.09 0.25 | 55.45 0.63 | 35.96 1.14 |
| Method | Drop Rate | |||||
| 0.1 | 0.3 | 0.5 | 0.7 | 0.8 | 0.9 | |
| Training-Dynamics-Based | ||||||
| D 2 | 7 1.42 | 6 9.90 | 6 7.68 | 6 5.01 | 61.58 | 52.79 |
| DUAL | 71.67 | 71.15 | 69.40 | 65.53 | 6 1.40 | 5 2.18 |
| Training-Dynamics-Free | ||||||
| Random | 71.32 | 70.24 | 68.49 | 64.31 | 59.96 | 50.75 |
Appendix figures & tables23 assets
Supplementary material from the paper’s appendix.
Appendix
| Method | Selection Requirement | Selection Strategy | KD Awareness | |||
| Dynamics- free | No extra optimization | Set-level selection | Budget- aware | KD- specific | Teacher relational geometry | |
| Forgetting ( Toneva et al., 2019 ) | ✗ | ✗ | ✗ | ✗ | ✗ | ✗ |
| EL2N ( Paul et al., 2021 ) | ✗ | ✗ | ✗ | ✗ | ✗ | ✗ |
| GraNd ( Paul et al., 2021 ) | ✗ | ✗ | ✗ | ✗ | ✗ | ✗ |
| CCS ( Zheng et al., 2023 ) | ✗ | ✗ | ✓ | ✓ | ✗ | ✗ |
| D 2 ( Maharana et al., 2024 ) | ✗ | ✗ | ✓ | ✓ | ✗ | ✗ |
| Drop Rate | Difficulty Region | Pool Size | |
| CIFAR-100 / Tiny-ImageNet | ImageNet | ||
| 0.3 | Medium | Hard | 80% |
| 0.5 | Medium | Medium | 60% |
| 0.7 | Medium | Medium | 60% |
| 0.8 | Easy | Medium | 40% |
| 0.9 | Easy | Medium-Easy | 20% |
| Relational Ordering | Score Geometry | Drop Rate | |
| 0.1 | 0.9 | ||
| Large | Large | 78.82 0.25 | 46.93 0.80 |
| Large | Small | 78.51 0.24 | 55.72 0.63 |
| Small | Large | 75.20 0.14 | 55.61 0.35 |
| Small | Small | 71.61 0.19 | 59.43 0.41 |
| Relational Ordering | Score Geometry | Drop Rate | |
| 0.1 | 0.9 | ||
| Large | Large | 78.82 0.25 | 46.93 0.80 |
| Large | Small | 78.94 0.21 | 50.38 0.71 |
| Small | Large | 78.69 0.25 | 47.43 0.58 |
| Small | Small | 78.12 0.15 ( ) | 54.10 1.19 |
| Relational Ordering | Score Geometry | Drop Rate | |
| 0.1 | 0.9 | ||
| Large | Large | 66.20 0.47 | 45.03 0.64 |
| Large | Small | 66.49 0.16 | 49.63 0.41 |
| Small | Large | 66.60 0.15 | 49.51 0.26 |
| Small | Small | 65.62 0.38 | 52.56 0.37 |
| Teacher W. | T. Train Acc. | T. Val. Acc. | Drop Rate | ||||||
| 0.0 | 0.1 | 0.3 | 0.5 | 0.7 | 0.8 | 0.9 | |||
| 128 | 99.98 | 80.95 | 79.35 | 78.49 | 75.87 | 72.50 | 66.20 | 60.82 | 44.81 |
| 64 (Default) | 99.97 | 7 9.71 | 7 9.73 | 78.82 | 76.69 | 73.16 | 67.07 | 61.35 | 46.93 |
| 32 | 99.98 | 77.87 | 79.80 | 7 8.80 | 7 7.16 | 7 4.40 | 68.88 | 63.81 | 47.84 |
| 16 | 99.30 | 73.17 | 78.65 | 78.12 | 77.17 | 75.08 | 71.31 | 6 7.49 | 54.10 |
| 8 | 91.87 | 67.44 | 75.42 | 75.33 | 74.64 | 73.54 | 7 1.24 | 69.09 | 62.49 |
| Teacher Width | T. Train Acc. | T. Val. Acc. | Drop Rate | ||||||
| 0.0 | 0.1 | 0.3 | 0.5 | 0.7 | 0.8 | 0.9 | |||
| 64 (Default) | 99.38 | 65.57 | 6 6.29 | 6 6.20 | 6 4.45 | 62.23 | 57.45 | 54.11 | 45.03 |
| 32 | 91.69 | 6 3.26 | 66.86 | 67.31 | 66.19 | 64.46 | 61.38 | 58.04 | 5 0.66 |
| 16 | 71.81 | 57.54 | 65.65 | 65.62 | 64.36 | 6 3.00 | 6 0.47 | 5 7.87 | 52.56 |
| 8 | 54.61 | 49.32 | 62.94 | 62.68 | 61.37 | 59.41 | 56.25 | 53.78 | 49.39 |
| Teacher | T. Train Acc. | T. Val. Acc. | Drop Rate | ||||||
| 0.0 | 0.1 | 0.3 | 0.5 | 0.7 | 0.8 | 0.9 | |||
| ResNet-50 | 92.93 | 80.11 | 71.43 | 71.08 | 69.85 | 67.75 | 63.18 | 58.40 | 49.03 |
| ResNet-34 | 85.06 | 76.32 | 71.68 | 71.32 | 70.24 | 68.49 | 64.31 | 59.96 | 50.75 |
| ResNet-18 | 78.62 | 71.47 | 71.43 | 71.01 | 70.05 | 68.45 | 64.99 | 61.05 | 52.66 |
| ResNet-10 | 75.77 | 68.23 | 70.73 | 70.38 | 69.48 | 68.03 | 64.85 | 61.40 | 54.04 |
| Teacher Embedding | T. Train Acc. | T. Val. Acc. | Drop Rate | ||||||
| 0.0 | 0.1 | 0.3 | 0.5 | 0.7 | 0.8 | 0.9 | |||
| 384 (Pre-trained) | 99.35 | 90.53 | 75.38 | 73.79 | 70.95 | 66.75 | 60.11 | 54.28 | 4 1.66 |
| 384 (From Scratch) | 99.97 | 68.58 | 70.41 | 69.31 | 66.08 | 61.64 | 55.13 | 49.59 | 37.79 |
| 192 | 9 9.96 | 6 8.81 | 70.82 | 68.85 | 64.61 | 57.75 | 49.13 | 40.39 | 28.44 |
| 96 | 95.06 | 65.74 | 7 2.69 | 7 0.66 | 66.95 | 60.34 | 51.70 | 43.54 | 30.45 |
| 48 | 65.85 | 49.76 | 71.06 | 69.84 | 6 7.64 | 6 4.38 | 5 8.06 | 5 4.01 | 43.68 |
| Teacher | KD Temperature | |||||
| Width-4 | 56.74 0.35 | 59.02 0.24 | 59.44 0.39 | 59.30 0.30 | 59.66 0.23 | 59.45 0.14 |
| Width-8 | 49.14 0.98 | 57.68 0.73 | 62.16 0.36 | 63.69 0.18 | 63.51 0.21 | 63.60 0.28 |
| Width-16 | 41.16 1.00 | 45.30 1.78 | 53.92 1.26 | 56.94 1.15 | 57.91 1.18 | 58.30 1.28 |
| Width-32 | 39.66 1.67 | 42.57 1.08 | 48.00 1.24 | 49.32 1.45 | 50.52 1.03 | 50.54 0.33 |
| Width-64 | 39.21 1.02 | 41.19 1.52 | 45.42 1.70 | 47.29 1.04 | 46.57 2.20 | 46.61 0.41 |
| Drop Rate | ||||||
| Selection Strategy | 0.1 | 0.3 | 0.5 | 0.7 | 0.8 | 0.9 |
| Filtered Random | 78.78 0.20 | 76.78 0.20 | 74.31 0.31 | 68.09 0.20 | 63.40 0.32 | 53.02 0.62 |
| Lowest Volume | 78.44 0.35 | 76.51 0.17 | 73.74 0.19 | 65.14 0.36 | 60.39 0.27 | 50.58 0.51 |
| Highest Volume (Ours) | 79.13 0.17 | 77.43 0.20 | 74.62 0.16 | 69.09 0.23 | 64.45 0.36 | 53.76 0.77 |
| Drop Rate | ||||||
| 0.1 | 0.3 | 0.5 | 0.7 | 0.8 | 0.9 | |
| 79.05 0.25 | 77.26 0.15 | 74.60 0.19 | 69.17 0.27 | 64.72 0.37 | 53.82 0.67 | |
| 79.02 0.29 | 77.39 0.20 | 74.72 0.22 | 68.95 0.28 | 64.62 0.22 | 54.19 0.73 | |
| 79.05 0.17 | 77.37 0.19 | 74.66 0.30 | 69.12 0.17 | 64.46 0.47 | 53.74 0.56 | |
| (default) | 79.13 0.17 | 77.43 0.20 | 74.62 0.16 | 69.09 0.23 | 64.45 0.36 | 53.76 0.77 |
| 79.04 0.27 | 77.56 0.28 | 74.73 0.21 | 68.85 0.30 | 64.42 0.26 | 53.47 0.53 | |
| Drop Rate | ||||||
| Proxy | 0.1 | 0.3 | 0.5 | 0.7 | 0.8 | 0.9 |
| Small (DVA, Ours) | 79.13 0.17 | 77.43 0.20 | 74.62 0.16 | 69.09 0.23 | 64.45 0.36 | 53.76 0.77 |
| Large @ ep10 | 79.02 0.24 | 76.86 0.24 | 73.53 0.26 | 67.27 0.23 | 62.28 0.25 | 52.47 0.44 |
| Large @ ep20 | 79.05 0.14 | 77.38 0.26 | 74.24 0.24 | 68.64 0.15 | 63.14 0.19 | 53.47 0.54 |
| Large @ ep30 | 79.26 0.28 | 77.14 0.19 | 74.24 0.15 | 68.70 0.31 | 63.33 0.20 | 52.63 0.56 |
| Large @ ep40 | 79.08 0.22 | 77.34 0.14 | 74.21 0.18 | 68.44 0.28 | 63.38 0.29 | 52.11 0.32 |
| Difficulty [-1.5pt]Filtering | Volume [-1.5pt]Maximization | Drop Rate | |||||
| 0.1 | 0.3 | 0.5 | 0.7 | 0.8 | 0.9 | ||
| Small | Small | 79.13 | 77.43 | 74.62 | 69.09 | 64.45 | 53.76 |
| Small | Random | 78.78 | 76.78 | 74.31 | 68.09 | 63.40 | 53.02 |
| Small | Large | 79.12 | 77.27 | 74.31 | 68.75 | 64.24 | 53.74 |
| Large | Small | 79.11 | 76.78 | 73.39 | 66.22 | 58.62 | 39.94 |
| Large | Large | 79.12 | 76.62 | 73.25 | 64.09 | 55.45 | 35.96 |
| Method | Drop Rate | |||||
| 0.1 | 0.3 | 0.5 | 0.7 | 0.8 | 0.9 | |
| Training-Dynamics-Based | ||||||
| D 2 | 7 3.56 0.24 | 71.14 1.05 | 66.60 0.88 | 60.24 0.40 | 55.30 0.12 | 44.75 1.11 |
| DUAL | 73.97 0.57 | 7 0.78 0.39 | 6 5.79 0.20 | 5 9.09 0.37 | 5 3.42 0.18 | 4 4.00 0.53 |
| Training-Dynamics-Free | ||||||
| Random | 7 3.79 0.72 | 70.95 0.23 | 66.75 0.40 | 6 0.11 0.41 | 5 4.28 0.67 | 4 1.66 0.87 |
| Method | Drop Rate | |||||
| 0.1 | 0.3 | 0.5 | 0.7 | 0.8 | 0.9 | |
| Training-Dynamics-Based | ||||||
| D 2 | 6 5.40 0.09 | 6 3.64 0.35 | 6 2.38 0.21 | 5 8.43 0.09 | 5 4.34 0.22 | 47.37 0.31 |
| DUAL | 66.32 0.25 | 65.26 0.16 | 63.49 0.31 | 59.49 0.42 | 55.32 0.24 | 4 6.30 0.17 |
| Training-Dynamics-Free | ||||||
| Random | 66.20 0.47 | 64.45 0.40 | 62.23 0.33 | 57.45 0.24 | 54.11 0.36 | 45.03 0.64 |
| Method | Drop Rate | |||||
| 0.1 | 0.3 | 0.5 | 0.7 | 0.8 | 0.9 | |
| Single-Teacher KD (Random) | ||||||
| Large T. | 78.82 0.25 | 76.69 0.17 | 73.16 0.20 | 67.07 0.49 | 61.35 0.43 | 46.93 0.80 |
| Small T. | 75.33 0.19 | 74.64 0.13 | 73.54 0.11 | 71.24 0.23 | 69.09 0.16 | 62.49 0.41 |
| Training-Dynamics-Based | ||||||
| Forgetting | 79.35 0.27 | 78.48 0.22 | 7 5.71 0.25 | 70.16 0.23 | 65.99 0.08 | 57.37 0.33 |
| Method | Drop Rate | |||||
| 0.1 | 0.3 | 0.5 | 0.7 | 0.8 | 0.9 | |
| Training-Dynamics-Based | ||||||
| D 2 | 7 3.56 0.24 | 70.87 0.74 | 65.97 0.47 | 59.50 0.13 | 54.89 0.20 | 45.96 0.31 |
| DUAL | 73.97 0.57 | 7 0.58 0.23 | 6 5.72 0.61 | 5 8.26 0.13 | 5 2.95 0.09 | 4 4.15 0.30 |
| Training-Dynamics-Free | ||||||
| Random | 7 3.79 0.72 | 7 0.99 0.63 | 65.97 0.27 | 58.90 0.32 | 5 4.01 0.47 | 44.05 0.44 |
| Method | Drop Rate | |||||
| 0.1 | 0.3 | 0.5 | 0.7 | 0.8 | 0.9 | |
| Training-Dynamics-Based | ||||||
| D 2 | 6 5.40 0.09 | 6 4.15 0.20 | 6 3.64 0.27 | 6 0.75 0.25 | 5 7.92 0.23 | 52.87 0.23 |
| DUAL | 66.32 0.25 | 65.84 0.19 | 64.68 0.36 | 61.34 0.03 | 58.03 0.30 | 5 2.69 0.09 |
| Training-Dynamics-Free | ||||||
| Random | 66.20 0.47 | 65.25 0.08 | 63.71 0.18 | 60.70 0.46 | 57.85 0.21 | 52.63 0.09 |
| Method | Drop Rate | |||||
| 0.1 | 0.3 | 0.5 | 0.7 | 0.8 | 0.9 | |
| Training-Dynamics-Based | ||||||
| D 2 | 7 3.09 0.26 | 7 1.58 0.27 | 6 8.80 0.40 | 6 2.70 0.24 | 5 9.49 0.09 | 4 9.34 0.66 |
| DUAL | 73.40 0.25 | 71.84 0.21 | 68.87 0.15 | 64.68 0.09 | 59.98 0.30 | 50.80 0.07 |
| Training-Dynamics-Free | ||||||
| Random | 7 3.00 0.07 | 70.22 0.12 | 6 6.10 0.44 | 59.89 0.11 | 54.90 0.37 | 43.40 1.20 |
| Method | Drop Rate | |||||
| 0.1 | 0.3 | 0.5 | 0.7 | 0.8 | 0.9 | |
| Training-Dynamics-Based | ||||||
| D 2 | 6 5.20 0.99 | 64.11 0.35 | 59.12 0.14 | 4 7.39 0.80 | 4 0.93 1.44 | 3 0.52 0.66 |
| DUAL | 66.15 0.79 | 6 3.93 0.28 | 5 8.69 0.37 | 49.74 0.99 | 42.64 0.78 | 31.89 0.72 |
| Training-Dynamics-Free | ||||||
| Random | 65.19 0.81 | 61.63 0.52 | 55.11 0.09 | 4 5.97 0.25 | 35.92 0.19 | 24.39 2.00 |
| Method | Selection Cost (PFLOPs) | Reduction |
| IF-Beta | 71.41 | - |
| DVA (Ours) | 2.81 |