cs.CVOct 6, 2026

REViT-v2: Hierarchical Windowed Roto-reflection Equivariant ViT for Equivariant Feature Extraction

Authors: Sheir A. Zaheer, Jihwan Moon, Chan Y. Park

Organizations: KC Machine Learning Lab, Rep. of Korea · NFOCZ Inc and Seoul National University, Rep. of Korea

Abstract

We propose a scalable roto-reflection-group-equivariant vision transformer based on windowed group-convolutional self-attention and a hierarchical feature architecture. We demonstrate that our approach can be scaled to group-equivariant vision transformers (ViTs) with millions of parameters and large datasets with practically sized images, i.e., ImageNet. The code and pretrained weights for the proposed Hierarchical Windowed Roto-reflection Equivariant ViTs (REViT-v2) are available at https://github.com/kc-ml2/revit.

Figures & tables

Appendix figures & tables1 asset

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

Jun 24, 2026cs.CV

REViT: Roto-reflection Equivariant Convolutional Vision Transformer

In this paper, we propose a discrete roto-reflection group equivariant vision transformer with convolutional attention. Roto-reflection equivariant networks preserve the rotational, flip and positional symmetry in feature maps, making them useful for tasks where orientation of the inputs is relevant to the model outputs. In image classification and object detection, most of the studies on roto-reflection equivariant models have focused on using convolutional neural networks rather than vision transformers. In this paper, we examine the challenges involved in achieving equivariance in vision transformers, and we propose a simpler way to implement a discretized roto-reflection group equivariant vision transformer. The experimental results demonstrate that our approach outperforms the existing approaches for developing discrete roto-reflection group equivariant neural networks for image classification.
Jun 26, 2026cs.CV

A Unified Framework for Vision Transformers Equivariant to Discrete Subgroups of O(2)\mathrm{O}(2)

Vision transformers have become a dominant architecture for visual recognition. However, standard models do not explicitly encode the planar symmetries that arise in many vision domains. We introduce a family of vision transformers equivariant to arbitrary discrete subgroups of O(2)\mathrm{O}(2), providing a unified framework that generalizes prior flipping- and D4D_4-equivariant transformer architectures. Our construction yields equivariant analogues of the core transformer components, together with expressivity guarantees for the resulting layers. In particular, we show that whenever H≤GH \le G, the class of GG-equivariant ViTs embeds naturally into the class of HH-equivariant ViTs. We also prove that, in the single-head setting, the corresponding equivariant self-attention layer realizes every GG-equivariant self-attention map representable by ordinary self-attention. We further construct a D6D_6-equivariant model based on hexagonal patches, making the architecture compatible with six-fold rotational symmetries. We evaluate the resulting models on the PatternNet aerial image dataset in artificially data-scarce regimes across subgroups of D4D_4 and D6D_6. Our experiments compare two equivariant attention mechanisms and analyze how the choice of homogeneous-space configurations used in the nonlinearities affects performance. Preliminary results under matched parameter budgets indicate that equivariance can improve recognition accuracy, motivating further study of how discrete symmetry groups shape transformer-based visual recognition models.
May 21, 2025cs.CV

Quick ViTs: Speeding up Vision Transformers through Equivariance

Natural images exhibit strong geometric regularities: local structures, such as edges, corners, and textures, appear in many orientations and mirror configurations. Since Vision Transformers (ViTs) operate on square image patches, these transformations naturally correspond to the dihedral symmetry group D8\mathrm{D}_8, also known as the octic group. Recent work has shown that ViTs can be made reflection equivariant and more efficient than standard ViTs simultaneously by implementing the linear layers in the Fourier domain of the reflection group. In this work, we extend the equivariance to reflections and rotations and analyze the scalability of the resulting networks. Our Quick ViTs, based on octic equivariant linear layers, achieve 5.33x reductions in FLOPs and up to 8x reductions in memory compared to ordinary linear layers. By analyzing the arithmetic intensity of these layers, we identify theoretical limits on how much the FLOP savings translate into throughput improvements on modern GPUs. However, these limitations disappear as the embedding dimensions increase. Enabled by their computational efficiency, we conduct a broader empirical evaluation of equivariant ViTs than in previous work. Upon training supervised (DeiT-III) and self-supervised (DINOv2) on ImageNet-1K, we find that our Quick ViTs match or exceed baseline accuracy while at the same time providing substantial efficiency gains.