In this work, we show that a single Transformer block, applied recurrently, can match the accuracy of a full-depth vision encoder at comparable inference FLOPs without intermediate feature distillation. reViT restores depth-specific transformations by representing the FFN at each recurrent depth as a convex combination of a small shared expert bank. A continuous normalized-depth coordinate programs this mixture, defining a resampleable trajectory through FFN parameter space. We evaluate this design in two regimes: supervised ImageNet-1k training and distillation from a DINOv2 teacher. Across both regimes, controlled adaptations identify weight-space merging as the strongest tested MoE family at a matching one-FFN budget, ahead of the token-dispatch and output-mixture alternatives. Trained from scratch, reViT-B/16 attains DeiT III accuracy with about 70% fewer stored parameters. An 8-experts model distilled using only the teacher's output features retains nearly all of its DINOv2 teacher's linear-probe accuracy and transfers across classification, segmentation, and depth prediction. Elastic-depth training allows one checkpoint (trained model) to operate at multiple tested depths by resampling the same normalized coordinate interval. For fixed-depth deployment, the recurrent block can be materialized as a conventional dense graph, removing online routing and merging without changing the one-FFN-per-depth compute but expanding deployment storage.
Figures & tables
Figure 1 : reViT overview. (a) One shared module runs for L steps. (b) At depth t , st softly mixes E expert parameter sets into one dense FFN shared by all tokens. (c) Selecting L resamples the depth program on a new grid. Teal marks shared parameters and amber marks depth and gating.
Scale
Method
k
E
Params (M)
GFLOPs/image
Top-1 (%) ↑
S/16
DeiT III ( Touvron et al., 2022 )
12
–
22.1
8.5
79.9
Raptor ( Jacobs et al., 2026 )
4
–
7.9
8.5
78.0
reViT (ours)
1
4
6.1
8.6
78.2
B/16
DeiT III
12
–
86.6
33.9
82.8
Raptor
4
–
29.9
33.9
82.8
reViT (ours)
1
4
23.6
34.3
83.0
Table 1 : ImageNet-1k training from scratch. All rows use the same recipe, and reViT values are means over three runs. Recurrent depths match the corresponding ViTs. k - counts distinct block templates and E - num. experts.
Method
Config
IN-1k
ADE20k
NYUv2 (lin)
NYUv2 (MLP)
Top-1 % ↑
mIoU % ↑
RMSE (m) ↓
RMSE (m) ↓
reViT (ours)
E=1
79.4
39.7
0.583
0.529
reViT (ours)
E=2
82.2
43.0
0.559
0.498
Raptor ( Jacobs et al., 2026 )
k=2
81.2
39.6
0.605
0.548
reViT (ours)
E=3
83.1
43.6
0.555
0.491
Raptor
k=3
83.0
43.0
0.567
0.506
Table 2 : DINOv2 distillation and frozen-probe transfer. reViT values are means over three runs. reViT uses final-layer supervision and Raptor intermediate features. Raptor IN-1k and ADE20k values are published, while NYUv2 values are our checkpoint reevaluations. E=k aligns stored FFN count, but not training. The E=8 row is compared only with DINOv2-B/14.
Top-1 (%) ↑ at nominal budget U
Family
Method
1×
1.5×
2×
4×
No MoE
reViT ( E=1 )
70.6
–
–
–
Weight merge
reViT (ours)
78.2
–
–
–
Lory ( Zhong et al., 2024 )
77.9
–
–
–
SMEAR ( Muqeeth et al., 2024 )
77.7
–
–
–
Token dispatch
DeepSeek-V3 ( Liu et al., 2024 )
70.6
71.3
72.4
76.5
Table 3 : MoE formulations under single-block recurrence: supervised ImageNet-1k training. All methods use the same -S backbone, L=12 , and base recipe, only the MoE mechanism changes. Higher- U settings also change expert count, activation count, and/or width.
Family
Method
IN-1k ↑
ADE20k ↑
NYUv2 lin. ↓
NYUv2 MLP ↓
No MoE
reViT ( E=1 )
79.4 / –
39.7 / –
0.583 / –
0.529 / –
Weight merge
reViT ( E=4 )
83.5 / –
44.6 / –
0.552 / –
0.491 / –
Lory ( Zhong et al., 2024 )
79.2 / –
39.9 / –
0.580 / –
0.531 / –
SMEAR ( Muqeeth et al., 2024 )
83.4 / –
44.3 / –
0.553 / –
0.485 / –
Token dispatch
DeepSeek-V3 ( Liu et al., 2024 )
74.4 / 76.6
36.6 / 39.9
0.654 / 0.607
0.582 / 0.539
Qwen3 ( Yang et al., 2025 )
– / 74.9
– / 40.1
– / 0.638
– / 0.550
Table 4 : MoE formulations under single-block recurrence: DINOv2 distillation. All entries use the same -B backbone, L=12 , and task loss. Each cell gives U=1/U=4 . Additional results are shown in Appendix B (Table 9 ).
Figure 2 : Routing variants under elastic-depth training. Both variants use the same per-image depth-dropout protocol and are evaluated at L∈{8,12,16,24} .
Number of experts E
Model
DeiT III
E=1
E=2
E=3
E=4
E=8
Δ1→8
reViT -S/16
79.9
70.6
74.2
76.5
78.2
80.2
+9.6
reViT -B/16
82.8
76.4
81.9
82.5
83.0
83.2
+6.8
reViT -L/16
84.1
78.3
83.0
83.5
83.8
84.0
+5.7
Table 5 : Effect of expert-bank size on ImageNet-1k. All models use supervised training.
Figure 8
Figure 5 : Available, selected, and realized expert directions. Effective ranks of the stacked expert W1 matrices, gate weights across depth, and stacked merged W1 matrices.
Appendix figures & tables10 assets
Supplementary material from the paper’s appendix.
Appendix
reViT -S/16
reViT -B/16
reViT -L/16
Optimization and schedule
Optimizer
AdamW
Peak learning rate
4×10−3
3×10−3
3×10−3
Effective batch size
2048
Schedule
cosine
Epochs
300
Appendix
Table 6 : Supervised ImageNet-1k configuration. The 300-epoch DeiT III base recipe is shared by all methods in Table 1 . Columns give scale-specific reViT settings, including recurrence and auxiliary losses.
Model
Params (M)
Latency (ms)
Throughput (img/s)
Memory (MiB)
B1 / B64
B1 / B64
B1 / B64
DeiT III-S/16
22.1
0.69 / 5.22
1,449 / 12,268
120.4 / 301.0
reViT -S/16 (dynamic)
6.1
0.78 / 7.26
1,282 / 8,814
63.1 / 292.4
Appendix
Table 7 : Compiled H100 inference. ViT-S models at 224×224 resolution in bfloat16. Deployment parameters, latency, throughput, and total runtime memory are shown for batches 1 and 64.
Method
U=1
U=1.5
U=2
U=4
Weight merge
reViT (ours)
4.7/18.9
–
–
–
Lory
4.7/18.9
–
–
–
SMEAR
4.7/18.9
–
–
–
Token dispatch
DeepSeek-V3
3.0/11.8
5.3/21.3
5.3/21.3
10.6/42.5
Appendix
Table 8 : Stored expert-FFN parameters (millions). Each entry gives the S/B counts.
Weight merge
Token dispatch
Output mixture
U
reViT ( E=4 )
Lory [1pt] ( Zhong et al., 2024 )
SMEAR [1pt] ( Muqeeth et al., 2024 )
DeepSeek-V3 [1pt] ( Liu et al., 2024 )
Qwen3 [1pt] ( Yang et al., 2025 )
MoEUT [1pt] ( Csordás et al., 2024 )
ReMoE [1pt] ( Wang et al., 2025 )
Soft-MoE [1pt] ( Puigcerver et al., 2024 )
ImageNet-1k linear top-1 (%, ↑ )
1×
83.5
79.2
83.4
74.4
–
75.1
79.6
78.5
1.5×
–
–
–
71.4
71.5
76.1
81.0
79.9
2×
–
–
–
75.9
73.6
75.2
81.5
80.7
4×
–
–
–
76.6
74.9
83.0
83.2
82.2
Appendix
Table 9 : Complete distilled recurrent-MoE comparison. Full results underlying Table 4 . All methods use the same B/14 recurrent backbone, L=12 , and task loss. U counts nominal expert-FFN arithmetic relative to one standard dense FFN per step. These fixed-depth results are separate from the elastic-depth evaluation in Fig. 2 . Higher- U settings also change the expert count, activation count, and/or width, as detailed below.
Method
Compact params (M)
GFLOPs
Top-1 (%) ↑
Depth index
online / folded
DeiT III-S/16
22.1
8.5 / 8.5
79.9
Block index
Tied ViT-S/16 + LoRA, r=64
7.3
10.4 / 8.5
77.6
Adapter index
reViT -S/16, E=4
6.1
8.6 / 8.5
78.2
Coordinate
Appendix
Table 10 : Depth-specific parameterizations on ImageNet-1k. Parameter counts refer to compact checkpoints. Online GFLOPs include parameter construction for reViT and low-rank updates for LoRA. Folded GFLOPs use 12 precomputed depth-specific dense weight sets. The reViT result is the mean of three runs. LoRA and DeiT III are single runs trained with the same supervised recipe.
Figure 6 : Depth-centered routing and merged-weight ranks. Centering removes the mean gate and merged W1 vectors across recurrent depth. The dashed line marks the resulting rank ceiling E−1 .
Figure 7 : Depthwise self-CKA after DINOv2 distillation. Linear CKA of mean-pooled patch-token residual states from the embedding through depth 12, computed over 1,000 ImageNet validation images, for DINOv2-B/14, Raptor- 4 , and distilled reViT -B/14 with E∈{1,8} and L=12 . The E=1 model is a plain tied block without an expert bank.
Figure 8 : Depthwise self-CKA after supervised ImageNet-1k training. Linear CKA of mean-pooled patch-token residual states over 1,000 validation images for DeiT III-S/16 and reViT -S/16 (12 blocks/steps), and DeiT III-L/16 and reViT -L/16 (24 blocks/steps). Dark near-diagonal bands indicate greater similarity between neighboring depths.
Figure 9 : Cross-CKA with DINOv2 across depth. Linear CKA between DINOv2-B/14 states (horizontal) and student states (vertical) for distilled reViT -B/14 with E∈{1,4,8} and Raptor- 4 , over the same 1,000-image subset as Figs. 7 and 8 . Dotted lines mark the maximum-CKA teacher depth, and dashed lines mark proportional depth. ρ is the Spearman correlation between recurrent depth and the maximum-CKA teacher depth. MAD is that path’s mean absolute deviation from proportional depth in teacher-layer units. Solid lines and bands show the weighted mean ± one standard deviation using weights proportional to max(CKA,0)4 .
Figure 10 : Schedule controls for expert ablation. (a) CKA change after ablation against each expert’s gate-weighted mean normalized depth. (b) Teacher CKA for the learned assignment (stars) and nonidentity permutations of gate columns across the fixed expert bank (dots). Horizontal lines show permutation means. We evaluate all 23 permutations for E=4 and 64 sampled permutations for E=8 .
Vision Transformers (ViTs) implement depth by stacking independently parameterized blocks, but it remains unclear how much of this parameterization is necessary and how much can be replaced by recurrent reuse. We study this question with bViT, a single-block recurrent ViT that repeatedly applies the same transformer block while preserving the iterative computation of a deep model. On ImageNet-1K, bViT-B reaches 0.779 validation accuracy compared with 0.789 for ViT-B under the same training recipe and computational budget, while using 8.6M rather than 86.6M parameters. This correspondence becomes stronger with model width, while narrow recurrent models exhibit a substantial performance gap. Beyond classification, the single-block formulation provides a controlled testbed for studying how transformer computation evolves with depth, since the same heads, neurons, and weight matrices can be tracked across recurrent steps. Analyses of attention, activation patterns, and step-conditioned spectral pruning reveal temporally organized behavior and step-dependent utilization of the shared parameters. bViT also transfers competitively to downstream tasks while enabling highly parameter-efficient adaptation. Our work shows that much of the performance associated with independently parameterized ViT depth can be recovered through recurrent reuse of a single sufficiently wide transformer block.
Michal Byra, Pawel Olszowiec, Grzegorz Stefanski +2
Samsung AI Center, Warsaw, Poland · Institute of Fundamental Technological Research, Polish Academy of Sciences, Warsaw, Poland
Vision Transformers (ViTs) achieve state-of-the-art segmentation accuracy but require large training datasets because each layer has unique parameters that must be learned independently. We present RD-ViT, a Recurrent-Depth Vision Transformer that adapts the Recurrent-Depth Transformer (RDT) architecture to dense prediction tasks, supporting both 2D and 3D inputs. RD-ViT replaces the deep stack of unique transformer blocks with a single shared block looped T times, augmented with LTI-stable state injection for guaranteed convergence, Adaptive Computation Time (ACT) for spatial compute allocation, depth-wise LoRA adaptation, and optional Mixture-of-Experts (MoE) feed-forward networks for category-specific specialization. We evaluate on the ACDC cardiac MRI segmentation benchmark in both 2D slice-level and 3D volumetric settings with exclusively real experiments executed in Google Colab. In 2D, RD-ViT outperforms standard ViT at 10% training data (Dice 0.774 vs 0.762) and at full data (0.882 vs 0.872). In 3D, RD-ViT with MoE achieves Dice 0.812 with 3.0M parameters, reaching 99.4% of standard ViT performance (0.817) at 53% of the parameter count. MoE expert utilization analysis reveals that different experts spontaneously specialize for different cardiac structures (RV, MYO, LV) without explicit routing supervision. ACT halting maps show higher compute allocation at cardiac boundaries, and the mean ponder time decreases from 2.6 to 1.4 iterations during training, demonstrating learned computational efficiency. Depth extrapolation enables inference with more loops than training without degradation. All code, notebooks, and results are publicly released.
Renjie He
Department of Radiation Oncology, The University of Texas MD Anderson Cancer Center, Houston, TX, USA
Vision Transformers (ViTs) achieve strong image-recognition performance, but their parameter count grows linearly with depth when each block is independently parameterized. Single-block recurrent ViTs (bViT) remove this growth by repeatedly applying one shared block. Rather than proposing a new architecture, we fix a bViT and provide a controlled empirical characterization of three training and inference regimes under a common CIFAR-100 protocol, asking: (i)~when does recurrence beat independently parameterized depth---at matched FLOPs or at matched parameter memory? (ii)~when a residual recurrent block is trained through an ODE solver, does solver order act as numerical refinement or as an architectural bias? and (iii)~what does robustness beyond the training horizon cost in nominal accuracy? We find that standard ViTs remain preferable when FLOPs are the primary constraint, whereas recurrent ViTs offer a better accuracy--parameter trade-off under memory constraints. Consistent with the standard view of residual networks as Euler discretizations of ODEs, the continuous-time analogue of a residual recurrent block is the state-subtracted vector field z˙=Fθ(z)−z; although known in principle, this distinction is easy to violate when the block is wrapped as a black-box vector field, and we qualify the cost at few accuracy points. Because the vector field is learned jointly with the solver, higher-order solvers act as a solver-induced architectural bias rather than a numerical-accuracy improvement, and their gains are not uniform. Finally, stage-wise deep supervision traces an accuracy--robustness frontier: it does not improve nominal accuracy, but degrades gracefully far beyond the training horizon, where naive recurrence collapses to near-random performance.
Grzegorz Gruszczynski, Pawel Olszowiec, Michal Byra +2
Samsung AI Center, Warsaw, Poland · Institute of Fundamental Technological Research, Polish Academy of Sciences, Warsaw, Poland