Are Coreset Selection Methods Worth Their Cost?
Organizations: Shandong University
Abstract
Coreset selection picks a representative subset of the labeled training set to make training cheaper. However, it is usually evaluated by downstream accuracy at a fixed subset size, ignoring both the time spent selecting the subset and the training recipe behind each reported number. We introduce an end-to-end benchmark that standardizes downstream training and charges selection and training to the same auditable wall-clock budget, spanning 4 datasets from CIFAR-10 to ImageNet-1K, 11 selectors, 5 fractions, and 3 seeds, with over 1,500 released runs. Repeated-sampling work has shown that budget-aware evaluation already favors random strategies. Our two budget studies test whether that verdict survives when every selector is granted its most favorable operating point. Across eight wall-clock budget anchors on each of CIFAR-10 and Tiny ImageNet, no anchor is won by a sophisticated selector: every winner is class-balanced random sampling, repeated random sampling, or full-data training. In fixed-budget duels on ImageNet-1K, training on all data for fewer epochs beats every selection strategy we probe while also costing the least. A per-dataset cost audit shows that selection cost is dominated at every scale by a fixed full-dataset scan, so it cannot be amortized away by selecting a smaller fraction, and its absolute size does not extrapolate from one dataset to another. We further quantify when selection does pay back through subset reuse, and document 9 correctness fixes to a widely used codebase, one of which shifts a standard Herding baseline by nearly 6 points. Selection time is not free preprocessing, and an evaluation that ignores it measures the wrong quantity.
Figures & tables
| Dataset | #Train | #Classes | Image size | Full-data acc. (%) |
|---|---|---|---|---|
| CIFAR-10 | 50,000 | 10 | ||
| Tiny ImageNet | 100,000 | 200 | ||
| CUB-200 | 5,994 | 200 | ||
| ImageNet-1K | 1,281,167 | 1000 |
| Method | ResNet-18, 200ep | ResNet-18, 50ep | ViT-Tiny, 300ep |
|---|---|---|---|
| Random | [3.33,4.00] | [3.78,5.00] | [3.94,5.28] |
| Uniform | [3.11,3.78] | [3.89,5.33] | [4.89,6.17] |
| Herding | [1.33,1.67] | [1.33,1.67] | [2.00,2.00] |
| Submodular | [4.56,5.44] | [3.11,4.44] | [3.78,4.89] |
| Craig | [10.33,10.67] | [10.67,10.67] | [6.72,7.11] |
| Forgetting | [2.33,2.78] | [1.50,2.17] | [1.00,1.00] |
Appendix figures & tables9 assets
Supplementary material from the paper’s appendix.
Appendix
| Location | Bug | Impact |
|---|---|---|
| herding.py:82 | Greedy selects farthest point ( argmax ) instead of nearest ( argmin ) | Accuracy ( pp on CIFAR-10) |
| submodular_fn.py:105 | GraphCut on wrong term of marginal gain | Accuracy |
| craig.py:48 | Dereferences non-existent dst_val.targets attribute | Feasibility (crash) |
| gradmatch.py:68 | Calls torch.lstsq , removed in PyTorch 2.x | Feasibility (crash) |
| gradmatch.py (OMP loop) | torch.cat in loop fragments CUDA memory | Feasibility (OOM on large classes) |
| resnet.py:105 | Hard-coded avg_pool2d(out, 4) breaks 64 64 inputs | Feasibility (shape error on Tiny ImageNet) |
| Dataset | Batch | Peak lr | Final lr | Epochs | Train-time input pipeline |
|---|---|---|---|---|---|
| CIFAR-10 | 256 | 200 | random crop 32 (pad 4), flip, normalize | ||
| Tiny ImageNet | 128 | 200 | random crop 64 (pad 8), flip, normalize | ||
| CUB-200 | 64 | 200 | random resized crop 224, flip, normalize | ||
| ImageNet-1K | 256 | 90 | random resized crop 224, flip, normalize |
| Anchor | Budget (s) | Winner | Score (%) | Runner-up | Margin (pp) | |
|---|---|---|---|---|---|---|
| 64 | Uniform | 85.21 | Submodular | 0.33 | ||
| q10 | 78 | Uniform | 85.21 | Submodular | 0.33 | |
| 132 | Full | 91.00 | RS2 | 0.57 | ||
| q50 | 351 | Full | 91.00 | GraNd | 0.20 | |
| 393 | Full | 94.18 | RS2 | 0.47 | ||
| 653 | Full | 94.82 | RS2 | 0.15 |
| Anchor | Budget (s) | Winner | Score (%) | Runner-up | Margin (pp) | |
|---|---|---|---|---|---|---|
| 451 | RS2 | 58.17 | Uniform | 1.16 | ||
| q10 | 482 | Full | 58.44 | RS2 | 0.09 | |
| 934 | Full | 63.92 | RS2 | 0.40 | ||
| q50 | 2302 | Full | 63.92 | RS2 | 0.13 | |
| 2769 | RS2 | 65.49 | Full | 0.10 | ||
| 4605 | RS2 | 66.02 | Full | 0.01 |
| Anchor | Peak winner | Peak acc. | Headline winner | Score (%) | Gap (pp) |
|---|---|---|---|---|---|
| CIFAR-10 | |||||
| C=10 | Uniform | 85.21 | Uniform | 85.21 | 0.33 |
| q10 | Uniform | 85.21 | Uniform | 85.21 | 0.33 |
| C=20 | Full | 91.06 | Full | 91.00 | 0.79 |
| q50 | Full | 91.06 | Full | 91.00 | 0.20 |
| C=60 | Full | 94.21 | Full | 94.18 | 0.71 |
| Group | Strategy | Sel. (h) | Train (h) | Total (h) | Top-1 (%) | |
|---|---|---|---|---|---|---|
| 18h | Full | (1,44) | 0.00 | 17.11 | 17.11 | |
| 18h | Uniform | (0.7,63) | 0.00 | 17.36 | 17.36 | |
| 18h | Herding | (0.5,59) | 5.45 | 11.83 | 17.28 | |
| 18h | Moderate | (0.5,62) | 5.01 | 12.43 | 17.44 | |
| 18h | kCenter | (0.5,62) | 5.01 | 12.43 | 17.44 | |
| 18h | Forgetting | (0.5,64) | 4.56 | 12.83 | 17.39 |
| CIFAR-10 | Tiny ImageNet | |||
|---|---|---|---|---|
| Method | ||||
| EL2N | 0.08 | 0.11 | 0.07 | 0.10 |
| Forgetting | 0.24 | 0.31 | 0.23 | 0.30 |
| Moderate | 0.23 | 0.29 | 0.22 | 0.28 |
| Herding | 0.24 | 0.31 | 0.23 | 0.30 |
| CIFAR-10 | Tiny ImageNet | |||
|---|---|---|---|---|
| Method | ||||
| Uniform | ||||
| Herding | ||||
| Moderate | ||||
| Forgetting | ||||
| EL2N | ||||
| Selection | ||
|---|---|---|
| EL2N, scored at epoch 1 | ||
| EL2N, scored at epoch 10 | ||
| EL2N, scored at epoch 20 | ||
| Random | ||
| Uniform |