The Composition Gap in Dataset Distillation
Organizations: Hokkaido University
Abstract
Dataset distillation compresses a training set into a small synthetic set, usually evaluated one at a time. In federated and data-governance settings, several parties distill their own data and a user trains on their union. We ask whether the union of separately distilled sets reproduces training on the union of the real data composability and show that it can fail even when every source is distilled exactly and the total budget admits an exact joint distillate. Compressing a training trajectory into fewer steps transforms the source statistics nonlinearly, so averaging compressed sources differs from compressing their average. For quadratic objectives we derive the exact composition error for two-to-one step compression in terms of the source-Hessian variance and the linear terms of the losses, and on a smooth network at small step sizes this prediction captures the local endpoint discrepancy in magnitude and direction. For learned synthetic sets, however, the composed error decomposes exactly into this local discrepancy and an aggregate source residual. Under endpoint matching the residual exceeds the structural term by more than an order of magnitude, and under distribution matching the two terms partly cancel. Joint distillation also retains an accuracy advantage when both sets are distilled from the same dataset, where the local discrepancy is exactly zero. Training fidelity and downstream accuracy are therefore distinct requirements, neither established by evaluating each set on its own.
Figures & tables
Appendix figures & tables32 assets
Supplementary material from the paper’s appendix.
Appendix
| Symbol | Meaning |
|---|---|
| Maps | |
| -step full-batch GD map at step size on dataset | |
| composed synthetic one-step map; real-union two-step map | |
| time-compression transform , | |
| Errors | |
| squared endpoint error over ; its RMS form | |
| rank of Gram | rel. Frobenius error | endpoint loss | measured/closed-form defect | log-log corr. | |
|---|---|---|---|---|---|
| 4 | 4 | 48.0% | 11.7 1.4 | 0.12 | |
| 6 | 6 | 31.7% | 7.6 0.9 | 0.16 | |
| 8 | 8 | 19.9% | 4.8 0.6 | 0.16 | |
| 12 | 12 | 1.000 0.000 | 1.00 | ||
| 24 | 12 | 1.000 0.000 | 1.00 |
| Component | Signature | Diagnostic | Minimum repair required |
|---|---|---|---|
| Global clock | common ratio | endpoint speed | scalar retiming |
| Source clock | coefficients | source HVP scale | per-source reweighting |
| Curvature variance | mode-dependent if the required retimings differ ( H.1 ) | ||
| Affine conflict | mean shift | scalar retiming only if it matches the linear part and the affine offset simultaneously (the scalar example of Section 4 is repaired by the factor ); otherwise mode-dependent | |
| Missing subspace | rank deficit | persistent residual | not repairable by retiming |
| Predictor | Spearman | partial (adj. label div.) | partial (adj. grad. dis.) |
|---|---|---|---|
| 0.991 [0.960, 0.999] | 0.869 [0.568, 0.989] | 0.701 [0.210, 0.971] | |
| 0.955 [0.857, 0.984] | 0.316 [ 0.316, 0.785] | 0.001 [ 0.262, 0.496] | |
| Label divergence | 0.964 [0.907, 0.985] | N/A | N/A |
| Gradient disagreement | 0.983 [0.937, 0.995] | N/A | N/A |
| Bank | Iter. | Bank loss | Held-out | ineq. | |||||
|---|---|---|---|---|---|---|---|---|---|
| 64 | 1,100 | 0.116 | 3.81 | 1.25 | 0.99 | 0.26 | 38.5 | 0/96 | |
| 64 | 3,000 | 0.114 | 3.79 | 1.25 | 0.99 | 0.26 | 38.4 | 0/96 | |
| 512 | 1,100 | 0.132 | 3.73 | 1.20 | 0.95 | 0.24 | 37.7 | 0/96 | |
| 512 | 3,000 | 0.130 | 3.71 | 1.20 | 0.95 | 0.25 | 37.5 | 0/96 |
| Level | Setting | IPC | rel. | ineq. | |||||||
| IID | , | 1 | 2.03 | 0.51 | 1.84 | 1.51 | 0.00009 | 0.00009 | 1.000 | 0.05 | 0/96 |
| IID | , | 5 | 1.10 | 0.27 | 1.04 | 1.01 | 0.00009 | 0.00009 | 1.000 | 0.04 | 0/96 |
| IID | , | 10 | 1.02 | 0.25 | 0.99 | 0.99 | 0.00009 | 0.00009 | 1.000 | 0.04 | 0/96 |
| IID | , | 5 | 4.11 | 0.26 | 3.86 | 3.71 | 0.0014 | 0.0014 | 0.997 | 0.03 | 0/96 |
| IID | , | 5 | 12.9 | 0.23 | 11.6 | 10.6 | 0.021 | 0.022 | 0.955 | 0/96 | |
| IID | , | 5 | 2.16 | 0.27 | 2.04 | 1.97 | 0.0005 | 0.0002 | 0.999 | 0.04 | 0/96 |
| Endpoint errors | Test accuracy | ||||||
| Source split | reference | separate | joint | excess | separate | joint | gain |
| Heterogeneity sweep (five partitions per level) | |||||||
| IID | .001) | .04) | .08) | .10) | .9) | .6) | .9) |
| Dir-1.0 | .99) | .29) | .08) | .25) | .6) | .5) | .6) |
| Dir-0.5 | .76) | .31) | .06) | .27) | .1) | .5) | .2) |
| Dir-0.1 | .17) | .33) | .26) | .44) | .7) | .1) | .7) |
| Dataset | Level | ||||||||
|---|---|---|---|---|---|---|---|---|---|
| CIFAR-10 | IID | 7.41 | 7.34 | 0.016 | 7.41 | 0.002 | 0.10 | ||
| CIFAR-10 | Dir-1.0 | 8.72 | 7.34 | 4.85 | 9.15 | 4.84 | 0.41 | 0.06 | |
| CIFAR-10 | Dir-0.5 | 8.71 | 7.33 | 5.58 | 9.84 | 7.26 | 0.51 | 0.01 | |
| CIFAR-10 | Dir-0.1 | 9.85 | 7.50 | 8.09 | 10.4 | 12.3 | 0.60 | 0.04 | |
| CIFAR-10 | no split | 7.43 | 7.34 | 0 | 7.43 | 0 | – | – | – |
| SVHN | IID | 6.08 | 6.13 | 0.013 | 6.08 | 0.002 | 0.08 | 0.09 |
| Dataset | Level | excess | ||||
|---|---|---|---|---|---|---|
| CIFAR-10 | IID | 7.41 / 13.4 | 7.34 / 13.4 | 7.41 / 13.4 | / | / |
| CIFAR-10 | Dir-1.0 | 8.72 / 13.4 | 7.34 / 13.4 | 9.15 / 13.6 | / | / |
| CIFAR-10 | Dir-0.5 | 8.71 / 13.5 | 7.33 / 13.3 | 9.84 / 13.9 | / | / |
| CIFAR-10 | Dir-0.1 | 9.85 / 13.7 | 7.50 / 13.4 | 10.4 / 14.4 | / | / |
| SVHN | Dir-0.1 | 7.65 / 11.0 | 6.18 / 10.7 | 12.1 / 12.8 | / | / |
| CIFAR-100 | Dir-0.1 | 2.51 / 7.83 | 1.96 / 7.82 | 2.87 / 7.85 | / | / |
| Dataset | Level | Weighting | removed (ind / joint) | excess | ||
|---|---|---|---|---|---|---|
| CIFAR-10 | IID | paper | 7.41 0.04 | 7.34 0.08 | – | 0.06 0.10 |
| bank | 7.82 0.05 | 7.83 0.08 | 0.11 / 0.14 | 0.01 0.12 | ||
| oracle | 5.83 0.05 | 5.76 0.04 | 0.38 / 0.38 | 0.07 0.04 | ||
| CIFAR-10 | Dir-1.0 | paper | 8.72 0.29 | 7.34 0.08 | – | 1.38 0.25 |
| bank | 7.93 0.08 | 7.83 0.08 | 0.17 / 0.14 | 0.10 0.08 | ||
| oracle | 6.00 0.10 | 5.76 0.04 | 0.53 / 0.38 | 0.24 0.10 |
| Set | Iterations | Accuracy | vs. union | effective rank |
|---|---|---|---|---|
| union of the independent sets | 0 | 46.44 0.50 | – | 8.91 |
| warm start from the union | 500 | 46.16 0.51 | 0.28 0.43 | 9.00 |
| warm start from the union | 2,000 | 46.88 0.67 | 0.44 0.47 | 9.15 |
| cold start from real images | 2,000 | 48.27 0.54 | 1.83 0.65 | 9.66 |
| joint set (official schedule) | 20,000 | 52.59 0.36 | 6.15 0.66 | 10.29 |
| Matched joint | Canonical joint | Matched | Canonical | |||||||
|---|---|---|---|---|---|---|---|---|---|---|
| Partition (images) | Acc. indep. | Acc. joint | Gain | Acc. joint | Gain | |||||
| 0 (150) | 9.48 | 7.92 | 1.55 | 7.95 | 1.53 | |||||
| 1 (170) | 8.88 | 7.82 | 1.06 | 7.35 | 1.53 | |||||
| 2 (200) | 12.18 | 7.29 | 4.90 | 7.29 | 4.90 | |||||
| IPC | acc. diff. (pp) | elasticity to next | ||||
|---|---|---|---|---|---|---|
| 1 | 0.416 0.023 | 21.3 | 2.60 0.45 | 4.47 0.07 | 0.66 1.12 | 0.83 |
| 5 | 0.206 0.021 | 11.9 | 1.44 0.10 | 2.71 1.70 | 5.25 1.50 | 0.88 |
| 10 | 0.172 0.019 | 10.2 | 1.22 0.03 | 2.68 2.00 | 5.47 0.90 | 1.09 |
| 20 | 0.157 0.019 | 9.20 | 1.09 0.06 | 2.61 2.10 | 2.06 0.14 | 0.95 |
| 50 | 0.147 0.017 | 8.63 | 1.02 0.08 | 2.62 2.30 | 1.68 0.27 | N/A |
| Dataset | Level | accuracy gain (pp) | |||
|---|---|---|---|---|---|
| CIFAR-10 | IID | 5 | 0.016 | 0.08 0.06 | 5.28 0.92 |
| CIFAR-10 | Dir-0.1 | 5 | 7.72 | 2.33 1.40 | 5.35 0.75 |
| SVHN | IID | 5 | 0.013 | 0.04 0.04 | 3.24 1.10 |
| SVHN | Dir-0.1 | 5 | 10.5 | 1.47 0.95 | 5.48 1.45 |
| CIFAR-10 | no split | 3 | 0 (exact) | 0.08 0.21 | 6.55 0.42 |
| SVHN | no split | 3 | 0 (exact) | 0.09 0.16 | 2.16 1.27 |
| ConvNet | ResNet-18 | ||||||
|---|---|---|---|---|---|---|---|
| Dataset | Level | joint | gain | union | joint | gain | |
| CIFAR-10 | IID | 5 | 51.6 0.6 | 5.28 0.92 | 40.3 0.6 | 42.1 0.5 | 1.81 0.85 |
| CIFAR-10 | Dir-1.0 | 5 | 51.9 0.5 | 5.72 0.64 | 40.2 0.6 | 42.3 0.6 | 2.05 0.77 |
| CIFAR-10 | Dir-0.5 | 5 | 51.7 0.5 | 5.86 1.24 | 39.2 1.1 | 42.6 0.4 | 3.31 0.84 |
| CIFAR-10 | Dir-0.1 | 5 | 51.1 1.1 | 5.35 0.75 | 37.8 1.1 | 41.2 0.8 | 3.44 0.95 |
| CIFAR-10 | no split | 3 | 52.6 0.4 | 6.55 0.42 | 41.1 0.8 | 42.2 0.1 | 1.13 0.79 |
| Level | eff. rank, separate (unioned) | eff. rank, joint | cross/within | |
|---|---|---|---|---|
| IID | 2 | 8.91 0.11 | 10.27 0.02 | 0.942 0.000 |
| Dirichlet-0.1 | 5 | 7.39 0.42 | 9.79 0.40 | 1.288 0.097 |
| No split (E7) | 3 | 8.91 0.05 | 10.29 0.11 | 0.941 0.001 |
| Network | Level | |||
|---|---|---|---|---|
| Local-law network at | ||||
| ReLU, 2,000-sample subset | IID / Dir-1.0 / Dir-0.1 | 0.112 / 0.116 / 0.126 | — | 0.29 0.08 |
| ReLU, subset | IID / Dir-0.1 | 0.096 / 0.108 | — | — |
| SiLU, subset | IID / Dir-0.1 | 0.047 / 0.066 | — | 0.14 0.02 |
| tanh, subset | IID / Dir-0.1 | 0.039 / 0.060 | — | 0.20 0.02 |
| Distillation network (ReLU) at | ||||
| Shift | label TV | † | |||||
|---|---|---|---|---|---|---|---|
| IID | 0.000 | 5 | 0.002 | 0.016 | (8.1) | 0.377 | |
| Feature skew f | 0.000 | 8 | 0.63 | 0.652 | (1.04) | 0.478 | |
| Dirichlet-1.0 | 0.227 | 5 | 4.61 | 4.54 | 1.01 | 0.928 | |
| Dirichlet-0.5 | 0.300 | 5 | 7.12 | 5.48 | 0.78 | not run | |
| Dirichlet-0.1 | 0.392 | 5 | 11.7 | 7.72 | 0.67 | 1.96 |
| ratio | |||
| Real partition and network only | |||
| ( ) | 8.09 1.17 | 14.1 2.3 | 1.75 |
| ( ) | 12.3 2.1 | 25.8 2.2 | 2.09 |
| Distilled sets (IPC 10 per source; total budget grows with ) | |||
| ( ) | 2.35 1.44 | 5.65 1.56 | 2.41 |
| accuracy gain (pp) | 5.35 0.75 | 7.34 1.34 | 1.37 |