One Block, Multiple Depths: Recurrent Vision Transformers with Depth-Programmed Experts
Organizations: Samsung AI Cambridge · Technical University of Iasi · Queen Mary University of London
Abstract
In this work, we show that a single Transformer block, applied recurrently, can match the accuracy of a full-depth vision encoder at comparable inference FLOPs without intermediate feature distillation. reViT restores depth-specific transformations by representing the FFN at each recurrent depth as a convex combination of a small shared expert bank. A continuous normalized-depth coordinate programs this mixture, defining a resampleable trajectory through FFN parameter space. We evaluate this design in two regimes: supervised ImageNet-1k training and distillation from a DINOv2 teacher. Across both regimes, controlled adaptations identify weight-space merging as the strongest tested MoE family at a matching one-FFN budget, ahead of the token-dispatch and output-mixture alternatives. Trained from scratch, reViT-B/16 attains DeiT III accuracy with about 70% fewer stored parameters. An 8-experts model distilled using only the teacher's output features retains nearly all of its DINOv2 teacher's linear-probe accuracy and transfers across classification, segmentation, and depth prediction. Elastic-depth training allows one checkpoint (trained model) to operate at multiple tested depths by resampling the same normalized coordinate interval. For fixed-depth deployment, the recurrent block can be materialized as a conventional dense graph, removing online routing and merging without changing the one-FFN-per-depth compute but expanding deployment storage.
Figures & tables
| Scale | Method | Params (M) | GFLOPs/image | Top-1 (%) | ||
|---|---|---|---|---|---|---|
| S/16 | DeiT III ( Touvron et al., 2022 ) | 12 | – | 22.1 | 8.5 | 79.9 |
| Raptor ( Jacobs et al., 2026 ) | 4 | – | 7.9 | 8.5 | 78.0 | |
| reViT (ours) | 1 | 4 | 6.1 | 8.6 | 78.2 | |
| B/16 | DeiT III | 12 | – | 86.6 | 33.9 | 82.8 |
| Raptor | 4 | – | 29.9 | 33.9 | 82.8 | |
| reViT (ours) | 1 | 4 | 23.6 | 34.3 | 83.0 |
| Method | Config | IN-1k | ADE20k | NYUv2 (lin) | NYUv2 (MLP) |
|---|---|---|---|---|---|
| Top-1 % | mIoU % | RMSE (m) | RMSE (m) | ||
| reViT (ours) | 79.4 | 39.7 | 0.583 | 0.529 | |
| reViT (ours) | 82.2 | 43.0 | 0.559 | 0.498 | |
| Raptor ( Jacobs et al., 2026 ) | 81.2 | 39.6 | 0.605 | 0.548 | |
| reViT (ours) | 83.1 | 43.6 | 0.555 | 0.491 | |
| Raptor | 83.0 | 43.0 | 0.567 | 0.506 |
| Top-1 (%) at nominal budget | |||||
| Family | Method | ||||
| No MoE | reViT ( ) | 70.6 | – | – | – |
| Weight merge | reViT (ours) | 78.2 | – | – | – |
| Lory ( Zhong et al., 2024 ) | 77.9 | – | – | – | |
| SMEAR ( Muqeeth et al., 2024 ) | 77.7 | – | – | – | |
| Token dispatch | DeepSeek-V3 ( Liu et al., 2024 ) | 70.6 | 71.3 | 72.4 | 76.5 |
| Family | Method | IN-1k | ADE20k | NYUv2 lin. | NYUv2 MLP |
|---|---|---|---|---|---|
| No MoE | reViT ( ) | 79.4 / – | 39.7 / – | 0.583 / – | 0.529 / – |
| Weight merge | reViT ( ) | 83.5 / – | 44.6 / – | 0.552 / – | 0.491 / – |
| Lory ( Zhong et al., 2024 ) | 79.2 / – | 39.9 / – | 0.580 / – | 0.531 / – | |
| SMEAR ( Muqeeth et al., 2024 ) | 83.4 / – | 44.3 / – | 0.553 / – | 0.485 / – | |
| Token dispatch | DeepSeek-V3 ( Liu et al., 2024 ) | 74.4 / 76.6 | 36.6 / 39.9 | 0.654 / 0.607 | 0.582 / 0.539 |
| Qwen3 ( Yang et al., 2025 ) | – / 74.9 | – / 40.1 | – / 0.638 | – / 0.550 |
| Number of experts | |||||||
|---|---|---|---|---|---|---|---|
| Model | DeiT III | ||||||
| reViT -S/16 | 79.9 | 70.6 | 74.2 | 76.5 | 78.2 | 80.2 | |
| reViT -B/16 | 82.8 | 76.4 | 81.9 | 82.5 | 83.0 | 83.2 | |
| reViT -L/16 | 84.1 | 78.3 | 83.0 | 83.5 | 83.8 | 84.0 | |
Appendix figures & tables10 assets
Supplementary material from the paper’s appendix.
Appendix
| reViT -S/16 | reViT -B/16 | reViT -L/16 | |
| Optimization and schedule | |||
| Optimizer | AdamW | ||
| Peak learning rate | |||
| Effective batch size | |||
| Schedule | cosine | ||
| Epochs | |||
| Model | Params (M) | Latency (ms) | Throughput (img/s) | Memory (MiB) |
|---|---|---|---|---|
| B1 / B64 | B1 / B64 | B1 / B64 | ||
| DeiT III-S/16 | 22.1 | 0.69 / 5.22 | 1,449 / 12,268 | 120.4 / 301.0 |
| reViT -S/16 (dynamic) | 6.1 | 0.78 / 7.26 | 1,282 / 8,814 | 63.1 / 292.4 |
| Method | ||||
|---|---|---|---|---|
| Weight merge | ||||
| reViT (ours) | 4.7/18.9 | – | – | – |
| Lory | 4.7/18.9 | – | – | – |
| SMEAR | 4.7/18.9 | – | – | – |
| Token dispatch | ||||
| DeepSeek-V3 | 3.0/11.8 | 5.3/21.3 | 5.3/21.3 | 10.6/42.5 |
| Weight merge | Token dispatch | Output mixture | ||||||
| reViT ( ) | Lory [1pt] ( Zhong et al., 2024 ) | SMEAR [1pt] ( Muqeeth et al., 2024 ) | DeepSeek-V3 [1pt] ( Liu et al., 2024 ) | Qwen3 [1pt] ( Yang et al., 2025 ) | MoEUT [1pt] ( Csordás et al., 2024 ) | ReMoE [1pt] ( Wang et al., 2025 ) | Soft-MoE [1pt] ( Puigcerver et al., 2024 ) | |
| ImageNet-1k linear top-1 (%, ) | ||||||||
| 83.5 | 79.2 | 83.4 | 74.4 | – | 75.1 | 79.6 | 78.5 | |
| – | – | – | 71.4 | 71.5 | 76.1 | 81.0 | 79.9 | |
| – | – | – | 75.9 | 73.6 | 75.2 | 81.5 | 80.7 | |
| – | – | – | 76.6 | 74.9 | 83.0 | 83.2 | 82.2 | |
| Method | Compact params (M) | GFLOPs | Top-1 (%) | Depth index |
|---|---|---|---|---|
| online / folded | ||||
| DeiT III-S/16 | 22.1 | 8.5 / 8.5 | 79.9 | Block index |
| Tied ViT-S/16 + LoRA, | 7.3 | 10.4 / 8.5 | 77.6 | Adapter index |
| reViT -S/16, | 6.1 | 8.6 / 8.5 | 78.2 | Coordinate |