cs.CVSep 14, 2025

Geometrically Constrained and Token-Based Probabilistic Spatial Transformers

Authors: Johann Schmidt, Tom Siegl, Martin Becker, Sebastian Stober

Organizations: Artificial Intelligence Lab, Otto-von-Guericke University Magdeburg, Germany · University of Rostock, Germany · Becker Lab, University of Marburg, Germany

Abstract

Spatial transformations such as rotation and scale obscure the morphological cues needed for accurate image classification. Careful consideration is required for reliable use in high stakes settings. A model should stay robust under such transformations, expose why a correction was applied, and signal when its input is ambiguous. While geometrically equivariant architectures provide a mathematically grounded solution, they often limit model flexibility through strict symmetry constraints and incur significant computational overhead. Spatial Transformer Networks (STNs) offer a data-driven, flexible alternative for learning pseudo-equivariances to affine transformations. However, STNs have historically been restricted to convolutional architectures and suffer from training instability. To address this, we introduce a novel STN framework. It leverages the global modeling capabilities of transformers to regress the affine transformation acting on the input. For this, we decompose affine transformations into interpretable primitives, regressed under adaptable geometric constraints, thereby preventing the training instability typically caused by degenerate transformations. By sharing weights between the localization network and the classification backbone, the framework requires minimal computational overhead. Extensive experiments on challenging insect biodiversity and medical imaging benchmarks demonstrate that our approach achieves superior predictive performance under diverse spatial transformations while maintaining high efficiency. Code is available at https://github.com/johSchm/TokenSTN.

Figures & tables

Appendix figures & tables1 asset

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. On Preserving Geometrical Invariance for Superpixel Image Classification using Graph Transformer

    Jul 5, 2026Sarabeshwar Balaji, Shubham Mohanty, Akash AnilCnn-Transformer TradeoffImage Classification

  2. Active Spatial Guidance: Eliminating Injected Positional Mechanisms in Vision Transformers

    Jul 1, 2026Cong Liu, Xiaofang Li, Simon X. YangVision TransformerVision Encoders

  3. INTCORT: Training-Free Spatial Reasoning Enhancement for Vision-Language Models via Input Transformations and Confidence Routing

    Sep 21, 2026Haoran Sun, Jingqi Xu, Yanhui Li +3Spatial ReasoningTransform