cs.CVOct 8, 2026

One Block, Multiple Depths: Recurrent Vision Transformers with Depth-Programmed Experts

Authors: Adrian Bulat, Yassine Ouali, Georgios Tzimiropoulos

Organizations: Samsung AI Cambridge · Technical University of Iasi · Queen Mary University of London

Abstract

In this work, we show that a single Transformer block, applied recurrently, can match the accuracy of a full-depth vision encoder at comparable inference FLOPs without intermediate feature distillation. reViT restores depth-specific transformations by representing the FFN at each recurrent depth as a convex combination of a small shared expert bank. A continuous normalized-depth coordinate programs this mixture, defining a resampleable trajectory through FFN parameter space. We evaluate this design in two regimes: supervised ImageNet-1k training and distillation from a DINOv2 teacher. Across both regimes, controlled adaptations identify weight-space merging as the strongest tested MoE family at a matching one-FFN budget, ahead of the token-dispatch and output-mixture alternatives. Trained from scratch, reViT-B/16 attains DeiT III accuracy with about 70% fewer stored parameters. An 8-experts model distilled using only the teacher's output features retains nearly all of its DINOv2 teacher's linear-probe accuracy and transfers across classification, segmentation, and depth prediction. Elastic-depth training allows one checkpoint (trained model) to operate at multiple tested depths by resampling the same normalized coordinate interval. For fixed-depth deployment, the recurrent block can be materialized as a conventional dense graph, removing online routing and merging without changing the one-FFN-per-depth compute but expanding deployment storage.

Figures & tables

Appendix figures & tables10 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Investigating Single-Block Recurrence in Vision Transformers for Image Recognition

    May 11, 2026Michal Byra, Pawel Olszowiec, Grzegorz Stefanski +2Vision TransformerRecurrent Transformers

  2. Training Crossroads for Recurrent Vision Transformers: Recurrence, Neural ODEs, and Deep Supervision

    Aug 5, 2026Grzegorz Gruszczynski, Pawel Olszowiec, Michal Byra +2Recurrent TransformersNeural ODEs