cs.LGOct 7, 2026

How Do Transformers Learn to Represent Symmetries?

Authors: Eduardo Santos-Escriche, Valerie Engelmayer, Ya-Wei Eileen Lin, Stefanie Jegelka

Organizations: Technical University of Munich Munich Center for Machine Learning · Massachusetts Institute of Technology

Abstract

Training Transformer-based architectures with finite data augmentation has become an increasingly popular approach in geometric machine learning. Despite its empirical success, the interplay between the Transformer architecture, invariance to different symmetries, and augmentation budgets remains underexplored. In this paper, we study the ability of a vanilla Transformer to learn various symmetries through finite data augmentation for point cloud datasets. We identify an ordering of increasing learnability across the following symmetry groups: (i) non-angle-preserving symmetries, (ii) angle-preserving symmetries, and (iii) base angle-preserving subgroups, such as translation, rotation, and scale. For the base angle-preserving groups, we further investigate the Transformer's extrapolation behavior and conduct a structural analysis of the trained models, allowing us to identify interpretable mechanisms that induce invariance. Finally, we extend our analysis to equivariant functions and show that the detected mechanisms for approximate invariance can also provide a key building block for learned equivariance. Our project page is available at https://transformers-learn-symmetries.github.io/

Figures & tables

Appendix figures & tables23 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Data Augmentation: A Fourier Analysis Perspective

    Jun 23, 2026Behrooz Tahmasebi, Melanie Weber, Stefanie JegelkaData AugmentationEquivariant Learning

  2. Rotational Equivariance in Machine Learning: A Comprehensive Tutorial

    Aug 31, 2026Peter Lippmann, Fred A. HamprechtRotation EquivarianceEquivariant Neural Networks