The standard Transformer architecture relies on a rigid pattern that alternates Attention and Feed-Forward Network (FFN) layers. Despite its widespread adoption, the inductive bias imposed by this strict separation has not been systematically examined. In this work, we investigate the necessity of the Attention-FFN dichotomy in Vision Transformers (ViTs). To facilitate this analysis, we introduce the AttenFeed module, a unified component that integrates the functional properties of both Attention and FFN. Based on this module, we devise the unified Vision Transformer (uViT), which replaces the conventional alternating Attention-FFN structure with a sequence of AttenFeed modules. We then use uViT as a control group that relaxes the Attention-FFN dichotomy of the standard ViT and systematically compare the two models across multiple datasets and model scales. Our experiments reveal that the Attention-FFN dichotomy can hinder performance at smaller model scales due to the rigid parameter allocation of ViTs. The AttenFeed module and uViT serve as new analytical tools for understanding the Attention-FFN structure and offer theoretical insights into the heuristically designed architecture of conventional ViTs.
Figures & tables
Figure 1: Architecture illustrations . (a) shows a layer of the standard Transformer, and (b) shows a layer of the unified Transformer, a control architecture for studying the Attention–FFN dichotomy.
Name
Params.
( L , H , D )
MLPs
ViT-S
21M
(12, 6, 384)
✓
uViT-S
20M
(12, 8, 640)
×
ViT-B
86M
(12, 12, 768)
✓
uViT-B
80M
(12, 16, 1280)
×
ViT-L
303M
(24, 16, 1024)
✓
uViT-L
289M
(28, 20, 1600)
×
Table 1: Configurations of the ViT and uViT models. L,H , and D denote the number of layers, heads, and model dimension, respectively.
CIFAR10
CIFAR100
Model
( L , H , D )
Params.
Test acc.
Rank(pre-head)
Rank(post-head)
Test acc.
Rank(pre-head)
Rank(post-head)
xViT
(6, 4, 128)
0.43M
78.34%
0.375
0.70
50.94%
0.406
0.25
uViT
84.00%
0.383
0.70
58.56%
0.508
0.28
xViT
(6, 4, 160)
0.67M
78.59%
0.356
0.70
52.73%
0.463
0.29
uViT
85.16%
0.369
0.70
59.80%
0.506
0.32
xViT
(12, 4, 128)
0.83M
80.13%
0.430
0.70
52.82%
0.586
0.29
Table 2: Performance and representation rank comparison on CIFAR10 and CIFAR100. The results demonstrate that uViT (composed of AttenFeed modules) consistently maintains a higher normalized rank and achieves superior accuracy compared to xViT (ViT without FFNs), confirming that the AttenFeed module effectively inherits the rank-preserving properties of FFNs.
ImageNet-Seg.
Pascal-VOC (Single class)
Model
Acc.
mIoU
mAP
Acc.
mIoU
mAP
ViT-S
0.772
0.629
0.821
0.745
0.550
0.853
uViT-S
0.699
0.538
0.806
0.689
0.490
0.847
ViT-B
0.725
0.569
0.801
0.707
0.480
0.844
uViT-B
0.680
0.516
0.775
0.682
0.449
0.828
ViT-L
0.666
0.500
0.765
0.677
0.422
0.819
Table 3: Segmentation performance using the attention maps from ViT and uViT. The results demonstrate that the AttenFeed module effectively generates meaningful segmentation masks, performing comparably to standard attention modules.
Figure 2: Qualitative comparison of attention maps between ViT and uViT.
Figure 3: Weight value distributions of uViT and standard ViT. The results show that uViT parameters occupy a transitional space between standard attention and FFN regimes, providing empirical evidence for the unified nature of the AttenFeed module.
Linear Probing (%)
Full Fine-tuning (%)
Model
SVHN
CIFAR10
CIFAR100
STL10
SVHN
CIFAR10
CIFAR100
STL10
ViT-S
40.03
89.94
80.84
58.64
93.03
75.68
51.09
52.13
uViT-S
56.57
95.40
90.99
72.67
91.76
73.77
47.77
50.63
ViT-B
47.91
91.73
73.10
96.93
91.08
73.58
44.73
50.04
uViT-B
60.84
93.75
77.85
97.36
91.70
74.98
47.41
51.56
ViT-L
50.69
93.23
76.05
97.63
87.35
70.10
43.90
47.66
Table 4: Transfer learning performance comparison between ViT and uViT. The results include Top-1 accuracy (%) for linear probing and full fine-tuning across four downstream datasets (SVHN, CIFAR10, CIFAR100, and STL10).
Test Acc. (%)
Model
ImageNet-1k
Places365
iNaturalist
ViT-S
53.64
44.63
52.82
uViT-S
67.11
50.25
66.81
ViT-B
74.22
53.37
76.28
uViT-B
74.90
54.45
78.26
ViT-L
77.33
54.75
85.39 †
Table 5: Pre-training performance across different model scales. uViT shows a significant performance advantage over ViT at the Small scale, while the gap narrows at the Base scale. † We doubled the batch size for the experiments, since convergence of uViT was unstable under the default setting.
Appendix figures & tables4 assets
Supplementary material from the paper’s appendix.
Appendix
SVHN
STL10
Model
( L , H , D )
Params.
Test acc.
Rank(pre-head)
Rank(post-head)
Test acc.
Rank(pre-head)
Rank(post-head)
xViT
(6, 4, 128)
0.43M
94.70%
0.320
0.8
67.15%
0.367
0.6
uViT
94.78%
0.320
0.8
69.79%
0.421
0.7
xViT
(6, 4, 160)
0.67M
94.87%
0.269
0.8
66.51%
0.325
0.7
uViT
95.00%
0.294
0.8
69.60%
0.406
0.7
xViT
(12, 4, 128)
0.83M
94.85%
0.336
0.8
67.31%
0.406
0.7
Appendix
Table 6: Performance and representation rank comparison on SVHN and STL10. uViT consistently matches or exceeds xViT in pre-head rank and test accuracy, confirming that the rank-preserving behavior of the AttenFeed module generalizes beyond CIFAR-10/100.
Figure 4: Qualitative comparison of attention maps between ViT and uViT. uViT-B reliably highlights the same salient object regions as ViT-B, supporting the attention-like behavior of the AttenFeed module.
Figure 5: Weight value distributions of uViT-S and ViT-S. The results show that uViT parameters occupy a transitional space between standard attention and FFN regimes at the Small scale, providing empirical evidence for the unified nature of the AttenFeed module.
Figure 6: Weight value distributions of uViT-L and ViT-L. The same transitional pattern holds at the Large scale, indicating that the unified nature of the AttenFeed module is not scale-specific.
A recent work shows that Attention Transfer, which transfers only the attention patterns from a pre-trained teacher Vision Transformer (ViT) to a randomly initialized standard student ViT, is sufficient to recover the full benefit of the teacher's pre-trained weights. We revisit this finding on a comprehensive benchmark of 20 teachers from 11 well-known ViT families and reveal that Attention Transfer is not universally effective. While 7 families transfer successfully, 4 consistently fail, falling up to 5.1% below the from-scratch no-transfer baseline. Further results demonstrate that this failure is family-consistent across model sizes, and persists under extended training durations, different transfer datasets, and out-of-distribution evaluations. Controlled analyses then consistently localize the problem to the attention-routing channel, indicating that the key issue is not whether the student can match the teacher's attention patterns, but whether the matched patterns remain functional for the student. Crucially, we identify architectural mismatch between the pre-trained teacher and the standard student as the primary mechanism. By adding only the teacher's native architectural components to the student in a randomly initialized state, we completely reverse the failure for all 4 families. Notably, these components alone do not improve from-scratch training, confirming that they specifically unlock the usability of the teacher's attention. We further systematically show that this failure is not explained by the inadequate choice of transfer loss or by differences in pre-training recipes. Our findings refine the prevailing understanding of attention in ViT representations: attention is sufficient \textit{only} when the student architecture matches the teacher.
Huaiyuan Qin, Muli Yang, Gabriel James Goenawan +4
Institute for Infocomm Research (I2R), A*STAR, Singapore · 3Shanghai Jiao Tong University · 2Sichuan University
The human visual system (HVS) employs foveated sampling and eye movements to achieve efficient perception, conserving both metabolic energy and computational resources. Drawing inspiration from this robustness and adaptability, we introduce the Foveated Dynamic Transformer (FDT), a foveation-guided dynamic token-selection architecture that integrates these mechanisms into a vision transformer framework. The FDT exhibits strong resilience to various types of noise and adversarial attacks, despite not being explicitly trained for such challenges. This inherent robustness is achieved through the use of fixation and foveation modules: the fixation module identifies fixation points to filter out irrelevant information, while the foveation module generates foveated embeddings with multi-scale information. At the 50% fixation-budget setting, FDT achieves higher accuracy than DeiT-S (81.9% vs. 80.9%) while reducing multiply-accumulate operations by 34.57%, highlighting one operating point on its accuracy-efficiency trade-off. These attributes position FDT as an HVS-inspired step toward artificial neural networks that combine adaptive computation with improved resilience.
Ibrahim Batuhan Akkaya, Kishaan Jeeveswaran, Bahram Zonooz +1
Advanced Research Lab, NavInfo Europe, Eindhoven, 5657 DB, Netherlands · Department of Mathematics and Computer Science, Eindhoven University of Technology, Eindhoven, 5612 AZ, Netherlands
Vision Transformers (ViTs) learn rich visual-semantic representations through all-to-all self-attention among patch tokens. However, this design implicitly assumes that direct pairwise patch interactions are necessary for effective representation learning. In this work, we challenge this assumption and show that representations supporting both global recognition and dense prediction can be learned without direct patch-to-patch interaction. We propose VECA (Visual Elastic-Core Attention), a vision transformer with core-periphery structured attention mediated by a small set of learned cores. Patch tokens exchange global information exclusively through these cores, while the full set of dense patches are preserved and iteratively updated across layers. This reduces attention complexity from O(N2) to O(N), linear in the number of patches N for a fixed core budget C. Unlike prior latent-token cross-attention architectures, VECA facilitates sparse global communication without compressing the spatial representation itself. Nested training along the core axis further enables a single model to elastically trade off computation and accuracy at inference time without retraining. Across image classification and dense prediction tasks, VECA remains competitive with full-attention backbones and outperforms the evaluated linear-complexity alternatives on most benchmarks. Moreover, without explicit supervision, these cores develop semantically organized structures that support object-label transfer across video frames. These results show that effective visual representations can be learned without direct all-to-all patch interaction.
Alan Z. Song, Yinjie Chen, Mu Nan +3
Carnegie Mellon University · University of Hong Kong