A Hypertoroidal Covering for Perfect Color Equivariance
Organizations: Princeton University · Tsinghua University
Abstract
When the color distribution of input images changes at inference, the performance of conventional neural network architectures drops considerably. A few researchers have begun to incorporate prior knowledge of color geometry in neural network design. These color equivariant architectures have modeled hue variation with 2D rotations, and saturation and luminance transformations as 1D translations. While this approach improves neural network robustness to color variations in a number of contexts, we find that approximating saturation and luminance (interval valued quantities) as 1D translations introduces appreciable artifacts. In this paper, we introduce a color equivariant architecture that is truly equivariant. Instead of approximating the interval with the real line, we lift values on the interval to values on the circle (a double-cover) and build equivariant representations there. Our approach resolves the approximation artifacts of previous methods, improves interpretability and generalizability, and achieves better predictive performance than conventional and equivariant baselines on tasks such as fine-grained classification and medical imaging tasks. Going beyond the context of color, we show that our proposed lifting can also extend to geometric transformations such as scale.
Figures & tables
| Network | A/A | A/B | A/C | Param |
|---|---|---|---|---|
| ResNet44 | 0.00 | 51.25 | 26.66 | 2.6M |
| CEConv-3 | 0.00 | 0.02 | 0.05 | 3.8M |
| CEConv-4 | 0.00 | 0.60 | 0.53 | 4.9M |
| LCER-H3 | 0.00 | 0.03 | 0.04 | 2.6M |
| LCER-H4 | 0.00 | 0.00 | 0.00 | 2.6M |
| CEN-H3 | 0.00 | 0.03 | 0.04 | 2.6M |
| Network | A/A | A/B | A/C | Param |
|---|---|---|---|---|
| ResNet44 | 0.00 | 41.40 | 42.20 | 2.6M |
| LCER-S3 | 0.00 | 0.00 | 0.04 | 2.6M |
| CEN-S3 | 0.00 | 0.00 | 0.00 | 2.6M |
| Network | A/A | A/B | A/C | Parameter |
|---|---|---|---|---|
| ResNet-18 | 8.32 | 37.70 | 33.88 | 11.2M |
| LCER-L3 | 7.31 | 34.83 | 36.43 | 11.1M |
| CEN-L3 | 5.36 | 14.42 | 31.53 | 11.2M |
| CEN-L4 | 6.60 | 13.09 | 29.53 | 11.2M |
| CEN-L8 | 5.67 | 11.88 | 27.46 | 11.1M |
| CEN-L16 | 5.66 | 11.09 | 24.32 | 11.2M |
| Network | Error | Param |
|---|---|---|
| ResNet44 | 55.40 (2.19) | 2.6M |
| LCER-H4S3 | 9.76 (3.54) | 2.6M |
| CEN-H3S3L3 | 3.01 (1.61) | 2.6M |
| CEN-H4S4L4 | 0.00 (0.00) | 2.6M |
| Network | Error | Param |
|---|---|---|
| ResNet50 | 28.91 (7.58) | 23.5M |
| CEConv-3 | 28.76 (9.93) | 23.1M |
| LCER-H4 | 27.53 (3.39) | 23.5M |
| LCER-S3 | 16.08 (2.68) | 23.3M |
| LCER-H4S3 | 19.06 (4.92) | 23.0M |
| CEN-H4 | 27.53 (3.39) | 23.5M |
| Network | Caltech 101 | CIFAR-10 | CIFAR-100 | Stanford Cars | Oxford Pets | STL-10 |
|---|---|---|---|---|---|---|
| Saturation Shifted Dataset | ||||||
| ResNet | 56.29 (0.59) | 11.71 (1.16) | 47.58 (5.19) | 37.34 (3.75) | 57.92 (2.49) | 30.28 (0.76) |
| ResNet-Jitter | 49.04 (0.63) | 11.91 (0.58) | 43.13 (2.13) | 32.00 (6.41) | 51.88 (1.05) | 28.86 (0.40) |
| ResNet-AugMix | 35.53 (2.50) | 11.80 (0.61) | 46.93 (0.95) | 30.21 (0.92) | 44.84 (1.11) | 23.57 (2.36) |
| ResNet-DeepAug | 42.87 (0.47) | 20.39 (1.44) | 51.64 (0.60) | 42.32 (1.11) | 43.57 (1.65) | 24.17 (1.49) |
| ResNet-Plackian | 38.05 (2.99) | 11.87 (1.24) | 47.12 (1.89) | 33.48 (3.16) | 41.43 (4.20) | 23.69 (2.79) |
Appendix figures & tables10 assets
Supplementary material from the paper’s appendix.
Appendix
| Original Dataset | ||||||
|---|---|---|---|---|---|---|
| Network | Caltech 101 | CIFAR-10 | CIFAR-100 | Stanford Cars | Oxford Pets | STL-10 |
| ResNet | 32.68 (1.55) | 7.86 (1.14) | 32.00 (0.63) | 25.41 (0.96) | 31.52 (2.05) | 18.59 (1.65) |
| ResNet-Gray | 33.79 (3.09) | 8.45 (0.68) | 32.04 (0.66) | 24.71 (0.93) | 30.38 (0.35) | 18.71 (1.47) |
| ResNet-Jitter | 32.90 (0.82) | 8.33 (0.44) | 32.27 (0.18) | 22.38 (1.65) | 30.06 (0.52) | 17.89 (1.48) |
| CEConv-3 | 34.74 (0.83) | 8.86 (0.33) | 34.95 (0.44) | 23.97 (1.56) | 31.08 (2.54) | 24.29 (1.31) |
| CEConv-4 | 33.52 (0.48) | 9.28 (0.24) | 35.46 (0.35) | 24.08 (0.66) | 33.70 (1.50) | 21.90 (1.64) |
| Caltech 101 | CIFAR-10 | CIFAR-100 | Stanford Cars | Oxford Pets | STL-10 | |
|---|---|---|---|---|---|---|
| Full | 0.80% | 0.57% | 0.69% | 0.88% | 0.65% | 0.89% |
| Partial | 0.69% | 0.64% | 0.64% | 0.67% | 0.48% | 0.66% |
| Caltech 101 | CIFAR-10 | CIFAR-100 | Stanford Cars | Oxford Pets | STL-10 | |
| Arch. | ResNet18 | ResNet44 | ResNet44 | ResNet18 | ResNet18 | ResNet18 |
| Batch Size | 16 | 128 | 128 | 16 | 16 | 16 |
| Epoch | 300 | 300 | 300 | 300 | 300 | 300 |
| Optimizer | Adam | SGD | SGD | Adam | Adam | Adam |
| Learning Rate | ||||||
| Scheduler | N/A | cos anneal. | cos anneal. | N/A | N/A | N/A |
| H1 | H4 | S4 | L4 | H4S4 | H4L4 | H4S4L4 | |
| Training Memory (MiB) | |||||||
| CEN | 1030 | 2739 | 2739 | 2739 | 9480 | 9480 | 32842 |
| LCER | 1014 | 2702 | 2702 | 2702 | 9906 | 9906 | not supported |
| Training Speed (it/s) | |||||||
| CEN | 19.38 | 11.33 | 11.34 | 11.27 | 5.23 | 5.24 | 1.59 |
| LCER | 19.01 | 11.30 | 11.31 | 11.31 | 5.22 | 5.21 | not supported |