cs.CVOct 6, 2026
SaveREViT-v2: Hierarchical Windowed Roto-reflection Equivariant ViT for Equivariant Feature Extraction
Organizations: KC Machine Learning Lab, Rep. of Korea · NFOCZ Inc and Seoul National University, Rep. of Korea
Abstract
We propose a scalable roto-reflection-group-equivariant vision transformer based on windowed group-convolutional self-attention and a hierarchical feature architecture. We demonstrate that our approach can be scaled to group-equivariant vision transformers (ViTs) with millions of parameters and large datasets with practically sized images, i.e., ImageNet. The code and pretrained weights for the proposed Hierarchical Windowed Roto-reflection Equivariant ViTs (REViT-v2) are available at https://github.com/kc-ml2/revit.
Figures & tables
| Model | Top-1 Accuracy (%) | Top-5 Accuracy (%) | Params. |
|---|---|---|---|
| ViT-S w/ aug | 72.08 | 89.54 | 22 M |
| RE-ResNet | 77.37 | 93.74 | 11 M |
| REViT-v2-T | 72.58 | 90.88 | 5 M |
| REViT-v2-S | 79.27 | 94.45 | 18 M |
| REViT-v2-S | 80.9 | 95.1 | 47 M |
Table 1: Performance Comparison on ImageNet-1K.
| Attention | Params. | Forward FLOPs | Latency vs global | Peak Training Memory | Accuracy (%) |
|---|---|---|---|---|---|
| Global | 97.7 K | 1.69 GFLOPs | 1 | 429.7 MB | 98.23 |
| Global w/downsampling | 97.7 K | 333.6 MFLOPs | 0.75 | 50.98 MB | 98.28 |
| Windowed | 102.9 K | 80.1 MFLOPs | 0.61 | 26.47 MB | 98.26 |
Table 2: Comparison with global G-CSA on Rotated MNIST.
Appendix figures & tables1 asset
Supplementary material from the paper’s appendix.
Appendix
| Group | Lifting Stem Equivariance Err. | Pre-Class Equivariance Err. |
|---|---|---|
| 0.000142 0.000017 | 0.00218 0.000471 | |
| 0.001316 0.000614 | 0.00008 0.000034 | |
| 0.001474 0.000366 | 0.000102 0.000013 |
Table 3: Equivariance error analysis