cs.CVOct 6, 2026
SaveREViT-v2: Hierarchical Windowed Roto-reflection Equivariant ViT for Equivariant Feature Extraction
Organizations: KC Machine Learning Lab, Rep. of Korea · NFOCZ Inc and Seoul National University, Rep. of Korea
Abstract
We propose a scalable roto-reflection-group-equivariant vision transformer based on windowed group-convolutional self-attention and a hierarchical feature architecture. We demonstrate that our approach can be scaled to group-equivariant vision transformers (ViTs) with millions of parameters and large datasets with practically sized images, i.e., ImageNet. The code and pretrained weights for the proposed Hierarchical Windowed Roto-reflection Equivariant ViTs (REViT-v2) are available at https://github.com/kc-ml2/revit.
Figures & tables
| Model | Top-1 Accuracy (%) | Top-5 Accuracy (%) | Params. |
|---|---|---|---|
| ViT-S w/ aug | 72.08 | 89.54 | 22 M |
| RE-ResNet | 77.37 | 93.74 | 11 M |
| REViT-v2-T | 72.58 | 90.88 | 5 M |
| REViT-v2-S | 79.27 | 94.45 | 18 M |
| REViT-v2-S | 80.9 | 95.1 | 47 M |
Table 1: Performance Comparison on ImageNet-1K.
| Attention | Params. | Forward FLOPs | Latency vs global | Peak Training Memory | Accuracy (%) |
|---|---|---|---|---|---|
| Global | 97.7 K | 1.69 GFLOPs | 1 | 429.7 MB | 98.23 |
| Global w/downsampling | 97.7 K | 333.6 MFLOPs | 0.75 | 50.98 MB | 98.28 |
| Windowed | 102.9 K | 80.1 MFLOPs | 0.61 | 26.47 MB | 98.26 |
Table 2: Comparison with global G-CSA on Rotated MNIST.
Appendix figures & tables1 asset
Supplementary material from the paper’s appendix.
Appendix
| Group | Lifting Stem Equivariance Err. | Pre-Class Equivariance Err. |
|---|---|---|
| 0.000142 0.000017 | 0.00218 0.000471 | |
| 0.001316 0.000614 | 0.00008 0.000034 | |
| 0.001474 0.000366 | 0.000102 0.000013 |
Table 3: Equivariance error analysis
Explore similar work
In this paper, we propose a discrete roto-reflection group equivariant vision transformer with convolutional attention. Roto-reflection equivariant networks preserve the rotational, flip and positional symmetry in feature maps, making them useful for tasks where orientation of the inputs is relevant to the model outputs. In image classification and object detection, most of the studies on roto-reflection equivariant models have focused on using convolutional neural networks rather than vision transformers. In this paper, we examine the challenges involved in achieving equivariance in vision transformers, and we propose a simpler way to implement a discretized roto-reflection group equivariant vision transformer. The experimental results demonstrate that our approach outperforms the existing approaches for developing discrete roto-reflection group equivariant neural networks for image classification.
A Unified Framework for Vision Transformers Equivariant to Discrete Subgroups of
Vision transformers have become a dominant architecture for visual recognition. However, standard models do not explicitly encode the planar symmetries that arise in many vision domains. We introduce a family of vision transformers equivariant to arbitrary discrete subgroups of , providing a unified framework that generalizes prior flipping- and -equivariant transformer architectures. Our construction yields equivariant analogues of the core transformer components, together with expressivity guarantees for the resulting layers. In particular, we show that whenever , the class of -equivariant ViTs embeds naturally into the class of -equivariant ViTs. We also prove that, in the single-head setting, the corresponding equivariant self-attention layer realizes every -equivariant self-attention map representable by ordinary self-attention. We further construct a -equivariant model based on hexagonal patches, making the architecture compatible with six-fold rotational symmetries. We evaluate the resulting models on the PatternNet aerial image dataset in artificially data-scarce regimes across subgroups of and . Our experiments compare two equivariant attention mechanisms and analyze how the choice of homogeneous-space configurations used in the nonlinearities affects performance. Preliminary results under matched parameter budgets indicate that equivariance can improve recognition accuracy, motivating further study of how discrete symmetry groups shape transformer-based visual recognition models.
Quick ViTs: Speeding up Vision Transformers through Equivariance
Natural images exhibit strong geometric regularities: local structures, such as edges, corners, and textures, appear in many orientations and mirror configurations. Since Vision Transformers (ViTs) operate on square image patches, these transformations naturally correspond to the dihedral symmetry group , also known as the octic group. Recent work has shown that ViTs can be made reflection equivariant and more efficient than standard ViTs simultaneously by implementing the linear layers in the Fourier domain of the reflection group. In this work, we extend the equivariance to reflections and rotations and analyze the scalability of the resulting networks. Our Quick ViTs, based on octic equivariant linear layers, achieve 5.33x reductions in FLOPs and up to 8x reductions in memory compared to ordinary linear layers. By analyzing the arithmetic intensity of these layers, we identify theoretical limits on how much the FLOP savings translate into throughput improvements on modern GPUs. However, these limitations disappear as the embedding dimensions increase. Enabled by their computational efficiency, we conduct a broader empirical evaluation of equivariant ViTs than in previous work. Upon training supervised (DeiT-III) and self-supervised (DINOv2) on ImageNet-1K, we find that our Quick ViTs match or exceed baseline accuracy while at the same time providing substantial efficiency gains.