cs.CVOct 6, 2026

REViT-v2: Hierarchical Windowed Roto-reflection Equivariant ViT for Equivariant Feature Extraction

Authors: Sheir A. Zaheer, Jihwan Moon, Chan Y. Park

Organizations: KC Machine Learning Lab, Rep. of Korea · NFOCZ Inc and Seoul National University, Rep. of Korea

Abstract

We propose a scalable roto-reflection-group-equivariant vision transformer based on windowed group-convolutional self-attention and a hierarchical feature architecture. We demonstrate that our approach can be scaled to group-equivariant vision transformers (ViTs) with millions of parameters and large datasets with practically sized images, i.e., ImageNet. The code and pretrained weights for the proposed Hierarchical Windowed Roto-reflection Equivariant ViTs (REViT-v2) are available at https://github.com/kc-ml2/revit.

Figures & tables

Appendix figures & tables1 asset

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. REViT: Roto-reflection Equivariant Convolutional Vision Transformer

    Jun 24, 2026Sheir A. Zaheer, Alexander C. Holston, Chan Y. ParkSelf-Supervised Vision TransformersRotation Equivariance

  2. Quick ViTs: Speeding up Vision Transformers through Equivariance

    May 21, 2025David Nordström, Johan Edstedt, Fredrik Kahl +1Self-Supervised Vision TransformersVision Transformer