Coreset selection picks a representative subset of the labeled training set to make training cheaper. However, it is usually evaluated by downstream accuracy at a fixed subset size, ignoring both the time spent selecting the subset and the training recipe behind each reported number. We introduce an end-to-end benchmark that standardizes downstream training and charges selection and training to the same auditable wall-clock budget, spanning 4 datasets from CIFAR-10 to ImageNet-1K, 11 selectors, 5 fractions, and 3 seeds, with over 1,500 released runs. Repeated-sampling work has shown that budget-aware evaluation already favors random strategies. Our two budget studies test whether that verdict survives when every selector is granted its most favorable operating point. Across eight wall-clock budget anchors on each of CIFAR-10 and Tiny ImageNet, no anchor is won by a sophisticated selector: every winner is class-balanced random sampling, repeated random sampling, or full-data training. In fixed-budget duels on ImageNet-1K, training on all data for fewer epochs beats every selection strategy we probe while also costing the least. A per-dataset cost audit shows that selection cost is dominated at every scale by a fixed full-dataset scan, so it cannot be amortized away by selecting a smaller fraction, and its absolute size does not extrapolate from one dataset to another. We further quantify when selection does pay back through subset reuse, and document 9 correctness fixes to a widely used codebase, one of which shifts a standard Herding baseline by nearly 6 points. Selection time is not free preprocessing, and an evaluation that ignores it measures the wrong quantity.
Figures & tables
Dataset
#Train
#Classes
Image size
Full-data acc. (%)
CIFAR-10
50,000
10
32×32
95.42
Tiny ImageNet
100,000
200
64×64
66.65
CUB-200
5,994
200
224×224
49.17
ImageNet-1K
1,281,167
1000
224×224
67.82
Table 1: Dataset summary and full-data baseline accuracy (ResNet-18 trained from scratch, mean over 3 seeds).
Method
ResNet-18, 200ep
ResNet-18, 50ep
ViT-Tiny, 300ep
Random
3.56 [3.33,4.00]
4.33 [3.78,5.00]
4.61 [3.94,5.28]
Uniform
3.44 [3.11,3.78]
4.67 [3.89,5.33]
5.50 [4.89,6.17]
Herding
1.56 [1.33,1.67]
1.50 [1.33,1.67]
2.00 [2.00,2.00]
Submodular
5.00 [4.56,5.44]
3.89 [3.11,4.44]
4.33 [3.78,4.89]
Craig
10.56 [10.33,10.67]
10.67 [10.67,10.67]
6.94 [6.72,7.11]
Forgetting
2.56 [2.33,2.78]
1.83 [1.50,2.17]
1.00 [1.00,1.00]
Table 2: Tiny ImageNet average ranks (lower is better) under the standard ResNet-18 200-epoch protocol and the ResNet-18 50-epoch and ViT-Tiny 300-epoch ablations. Brackets give 95% bootstrap intervals over seed resamples. Bold marks the rank-1 method in each column.
Appendix figures & tables9 assets
Supplementary material from the paper’s appendix.
Appendix
Location
Bug
Impact
herding.py:82
Greedy selects farthest point ( argmax ) instead of nearest ( argmin )
Table 3: Fixes applied to the inherited DeepCore code. Impact marks accuracy-changing versus feasibility issues.
Dataset
Batch
Peak lr
Final lr
Epochs
Train-time input pipeline
CIFAR-10
256
0.1
10−4
200
random crop 32 (pad 4), flip, normalize
Tiny ImageNet
128
0.1
10−4
200
random crop 64 (pad 8), flip, normalize
CUB-200
64
0.01
10−6
200
random resized crop 224, flip, normalize
ImageNet-1K
256
0.1
10−4
90
random resized crop 224, flip, normalize
Appendix
Table 4: Downstream training protocol per dataset (fixed across all methods and fractions).
Anchor
Budget (s)
Winner
(f,E)
Score (%)
Runner-up
Margin (pp)
C=10
64
Uniform
(0.7,14)
85.21
Submodular
0.33
q10
78
Uniform
(0.7,14)
85.21
Submodular
0.33
C=20
132
Full
(1.0,20)
91.00
RS2
0.57
q50
351
Full
(1.0,20)
91.00
GraNd
0.20
C=60
393
Full
(1.0,60)
94.18
RS2
0.47
C=100
653
Full
(1.0,100)
94.82
RS2
0.15
Appendix
Table 5: CIFAR-10 anchors, priced on the single-machine timing audit: winner and runner-up by three-seed mean headline score among affordable cells. Scores use recorded final accuracy, with the favorable-bound treatment for missing sophisticated canonical finals. The single-seed lr 0.05 ablation reaches a peak of 87.77 at the q10 budget but is not eligible (Appendix H ). Submodular is priced at zero selection cost.
Anchor
Budget (s)
Winner
(f,E)
Score (%)
Runner-up
Margin (pp)
C=10
451
RS2
(0.7,14)
58.17
Uniform
1.16
q10
482
Full
(1.0,10)
58.44
RS2
0.09
C=20
934
Full
(1.0,20)
63.92
RS2
0.40
q50
2302
Full
(1.0,20)
63.92
RS2
0.13
C=60
2769
RS2
(0.7,86)
65.49
Full
0.10
C=100
4605
RS2
(0.7,143)
66.02
Full
0.01
Appendix
Table 6: Tiny ImageNet anchors, priced on the single-machine timing audit, using the same final/favorable-bound score as Table 5 . All eight go to RS2 or full-data training, with the other of the two as runner-up at seven of eight. Uniform is the runner-up at the cheapest anchor.
Anchor
Peak winner
Peak acc.
Headline winner
Score (%)
Gap (pp)
CIFAR-10
C=10
Uniform
85.21
Uniform
85.21
0.33
q10
Uniform
85.21
Uniform
85.21
0.33
C=20
Full
91.06
Full
91.00
0.79
q50
Full
91.06
Full
91.00
0.20
C=60
Full
94.21
Full
94.18
0.71
Appendix
Table 7: Historical peak scores versus the headline recorded-final/favorable-bound scores at the same audited anchors. Gap is the headline winner minus the best affordable sophisticated method under the headline rule.
Group
Strategy
(f,E)
Sel. (h)
Train (h)
Total (h)
Top-1 (%)
18h
Full
(1,44)
0.00
17.11
17.11
66.17±0.07
18h
Uniform
(0.7,63)
0.00
17.36
17.36
65.83±0.05
18h
Herding
(0.5,59)
5.45
11.83
17.28
64.12±0.06
18h
Moderate
(0.5,62)
5.01
12.43
17.44
63.98†
18h
kCenter
(0.5,62)
5.01
12.43
17.44
63.87†
18h
Forgetting
(0.5,64)
4.56
12.83
17.39
63.24†
Appendix
Table 8: ImageNet operating points priced from the timing probes. Final-epoch accuracy is mean ± population standard deviation over three seeds, and daggers mark one seed. Uniform’s nonzero indexing cost rounds to 0.00h.
CIFAR-10
Tiny ImageNet
Method
f=0.1
f=0.3
f=0.1
f=0.3
EL2N
0.08
0.11
0.07
0.10
Forgetting
0.24
0.31
0.23
0.30
Moderate
0.23
0.29
0.22
0.28
Herding
0.24
0.31
0.23
0.30
Appendix
Table 9: R18 cost crossover K∗ versus full-data R18 training at 200 epochs. Actual reuse counts are integers.
CIFAR-10
Tiny ImageNet
Method
f=0.1
f=0.3
f=0.1
f=0.3
Uniform
+2.2
+1.5
+1.8
+0.0
Herding
+2.0
+1.5
+5.1
+2.5
Moderate
−0.8
+0.7
+2.5
−0.7
Forgetting
−26.1
−0.0
+1.1
+3.1
EL2N
−46.4
−6.8
−19.1
−10.8
Appendix
Table 10: Separate R18-to-R50 transfer experiment: accuracy difference (pp) from R50 trained on same-size random subsets.
Selection
f=0.05
f=0.1
EL2N, scored at epoch 1
14.12±1.46
28.80±5.95
EL2N, scored at epoch 10
25.07±0.13
38.23±1.72
EL2N, scored at epoch 20
32.13±1.35
46.44±1.55
Random
61.96±0.90
72.76±0.77
Uniform
60.72±2.71
72.66±1.50
Appendix
Table 11: CIFAR-10 EL2N with later scoring epochs: final test accuracy, mean ± sample SD over three seeds, controls rerun in the same batch.