Copy the Same, Distill the Difference: Initializing Linear Vision Transformers
Authors: Huaiyuan Qin, Muli Yang, Gabriel James Goenawan, Shiqi Huang, Min Kass Chong, Wahyu Wiratama, Peng Hu, Chen Gong, +4 more
Organizations: Institute of Advanced Intelligence and Computing (IAIC), A*STAR, Singapore · Nanyang Technological University · ST Engineering Geo-Insights, Singapore · Sichuan University · Shanghai Jiao Tong University · University of Science and Technology of China
Linear Vision Transformers (ViTs) are designed to replace the attention in Softmax ViTs with the linear-complexity attention operator for more efficient token routing, but they require from-scratch pre-training and typically underperform the original Softmax version. How to initialize linear ViTs both efficiently and effectively still remains unclear. In this work, we explicitly ask: given that most foundation ViTs are built on the mainstream Softmax attention, can linear ViTs benefit from their pre-trained weights? Recent works on Attention Transfer show that attention is the effective transferable component between Softmax ViTs, suggesting attention alone suffices for such reuse. However, we find the opposite for Softmax-to-linear transfer. The attention weights are operator-specific: copying them barely helps, and is sometimes even worse than random initialization. Instead, the attention's token routing behavior can be recovered through distillation with a proper loss design, letting linear ViTs reduce the gap and even match Softmax ones. In contrast, the MLP weights, which carry the learned representation, are operator-agnostic: they can be transferred by simple direct copying, which already carries most of the benefit of the pre-trained weights. Thus, copying MLPs can serve as an effective foundation for Softmax-to-linear transfer: paired with the distilled attention, linear ViTs eventually close the remaining gap and even surpass Softmax ones. These findings hold consistently across various linear ViT variants, different model sizes, and diverse datasets. We hope this study deepens the understanding of reusing pre-trained weights across attention operators: copy what stays the same and distill what differs, to recover the benefit across the Softmax-to-linear boundary.
Figures & tables
Figure 1: What survives the Softmax-to-linear boundary? We ask whether linear ViTs can benefit from pre-trained Softmax weights for initialization. We find that the attention weights are operator-specific that copying them barely helps. Their token routing behavior can instead be recovered by distilling the attention outputs, while the operator-agnostic MLP weights, which carry the learned representation, can be simply transferred by direct copying. Paired together, linear ViTs can match and even surpass the Softmax ones, consistently across diverse settings.
ELU
ReLU
MHLA
TTT
Random
79.6
76.5
81.8
82.2
DeiT
79.8 +0.2
75.6 -0.9
81.7 -0.1
81.0 -1.2
DINO
80.2 +0.6
75.2 -1.3
81.6 -0.2
80.6 -1.6
MoCov3
79.9 +0.3
75.7 -0.8
81.7 -0.1
81.1 -1.1
iBOT
80.3 +0.7
75.3 -1.2
81.8 +0.0
80.7 -1.5
MAE
80.4 +0.8
76.0 -0.5
82.0 +0.2
81.4 -0.8
Table 1: Pre-trained weights robustness. We repeat the attention-only transfer on ViT-Base with different pre-trained Softmax weights, under the same ImageNet-1K evaluation. Gray row shows the shared Random baseline, and colored deltas indicate gains/drops relative to it. The failure remains consistent across all pre-trained sources: none provides a reliable advantage over random initialization.
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
Setting
Copied components
Class token
Pos. embedding
Random
–
random
random
Mimetic
–
random
random
Impulse
–
random
random
Copy
Attention
random
random
MLPCopy
MLPs
random
random
FullCopy
all Softmax weights
copied
copied
Appendix
Table A: Implementation details of transfer protocol. We list the weight-initialization details of each component in the Softmax-to-linear transfer experiments. Note that we follow the default implementations in MHLA and TTT to use sinusoidal positional embeddings and average pooling with no class token. All other settings keep the class token and learned positional embeddings. Mimetic and Impulse assume the standard QK⊤ attention parameterization and are therefore incompatible with TTT.
Dataset
Classes
Train
Test / Val
Epochs
Warm-up
ImageNet-1K ( Deng et al., 2009 )
1,000
1,281,167
50,000
300
50
CIFAR-10 ( Krizhevsky, 2009 )
10
50,000
10,000
300
50
CIFAR-100 ( Krizhevsky, 2009 )
100
50,000
10,000
300
50
STL-10 ( Coates et al., 2011 )
10
5,000
8,000
300
50
Food ( Bossard et al., 2014 )
101
75,750
25,250
300
50
Flowers ( Nilsback and Zisserman, 2008 )
102
1,020
6,149
600
100
Appendix
Table B: Details of datasets and training schedules.
Config
Transfer datasets
ImageNet-1K
Optimizer
AdamW
AdamW
Base Learning Rate
2e-3
2e-3
Weight Decay
0.05
0.05
Optimizer Momentum
β1=0.9 , β2=0.999
β1=0.9 , β2=0.999
Layer-wise LR Decay
–
–
Batch Size
512
4096
Appendix
Table C: Training recipes for transfer datasets and ImageNet-1K.
w/
IN-1K
COCO
ADE20K
Distill
Acc
AP b
AP 50b
AP 75b
AP m
AP 50m
AP 75m
mIoU
+MS
Softmax
81.8
42.9
65.7
46.8
39.4
62.6
42.0
46.1
47.1
Random
76.5
34.0
56.5
37.6
32.1
55.5
34.4
34.9
35.7
✓
79.4 +2.9
36.2 +2.2
59.0 +2.5
40.3 +2.7
35.8 +3.7
59.2 +3.7
37.9 +3.5
40.7 +5.8
41.8 +6.1
Copy
75.6
33.3
55.9
37.1
31.6
55.0
33.7
35.2
35.8
✓
79.9 +4.3
36.9 +3.6
59.3 +3.4
40.7 +3.6
36.0 +4.4
59.6 +4.6
38.2 +4.5
41.5 +6.3
42.1 +6.3
Appendix
Table D: ImageNet-1K and dense prediction results with ReLU at Base size. All settings adopt transferred weights from DeiT and are then fine-tuned with Mask R-CNN (1 × ) on COCO and UperNet (160k iters) on ADE20K. We report the Top-1 Accuracy (%) on ImageNet-1K, box and mask AP on COCO, and mIoU with single and multi-scale (+MS) on ADE20K, respectively. w/ Distill with ✓ indicates applying attention distillation during the Softmax-to-linear transfer, and colored deltas indicate gains relative to the same setting w/o distillation. Softmax row refers to the Softmax reference and the plain-ViT baselines of Chen et al. (2023) .
Weight Copying
Accuracy
Attn.
MLP
Patch emb.
Pos. emb.
Cls. token
w/o Distill
w/ Distill
FullCopy
✓
✓
✓
✓
✓
80.8 +4.3
82.0 +2.6
Variant 1
✓
✓
✓
✓
80.7 +4.2
81.8 +2.4
Variant 2
✓
✓
✓
80.6 +4.1
81.9 +2.5
Variant 3
✓
✓
80.7 +4.2
81.9 +2.5
MLPCopy
✓
80.2 +3.7
81.4 +2.0
Appendix
Table G: Ablation on copying the remaining components. Starting from FullCopy , we gradually remove the weight loading of the class token, positional embedding, and patch embedding, while all attention and MLP weights remain copied. We report the Top-1 Accuracy (%) on ImageNet-1K with ReLU at Base size, with colored deltas relative to the Random baseline.
While linear-complexity attention mechanisms offer a promising alternative to Softmax attention for overcoming the quadratic bottleneck, training such models from scratch remains prohibitively expensive. Inheriting weights from pretrained Transformers provides an appealing shortcut, yet the fundamental representational gap between Softmax and linear attention prevents effective weight transfer. In this work, we address this conversion challenge from two perspectives: architectural alignment and representational alignment. We identify Test-Time Training (TTT) as a linear-complexity architecture whose two-layer dynamic formulation is structurally aligned with Softmax attention, enabling direct inheritance of pretrained attention weights. To further align representational properties, including key shift-invariance and locality, we introduce key instance normalization and a lightweight locality enhancement module. We validate our approach by linearizing Stable Diffusion 3.5 and introduce SD3.5-T5 (Transformer To Test Time Training). With only 1 hour of fine-tuning on 4×H20 GPUs, SD3.5-T5 achieves comparable text-to-image quality to the fine-tuned Softmax model, while accelerating inference by 1.32× and 1.47× at 1K and 2K resolutions. Code is available at https://github.com/LeapLabTHU/Transformer-to-TTT.
A recent work shows that Attention Transfer, which transfers only the attention patterns from a pre-trained teacher Vision Transformer (ViT) to a randomly initialized standard student ViT, is sufficient to recover the full benefit of the teacher's pre-trained weights. We revisit this finding on a comprehensive benchmark of 20 teachers from 11 well-known ViT families and reveal that Attention Transfer is not universally effective. While 7 families transfer successfully, 4 consistently fail, falling up to 5.1% below the from-scratch no-transfer baseline. Further results demonstrate that this failure is family-consistent across model sizes, and persists under extended training durations, different transfer datasets, and out-of-distribution evaluations. Controlled analyses then consistently localize the problem to the attention-routing channel, indicating that the key issue is not whether the student can match the teacher's attention patterns, but whether the matched patterns remain functional for the student. Crucially, we identify architectural mismatch between the pre-trained teacher and the standard student as the primary mechanism. By adding only the teacher's native architectural components to the student in a randomly initialized state, we completely reverse the failure for all 4 families. Notably, these components alone do not improve from-scratch training, confirming that they specifically unlock the usability of the teacher's attention. We further systematically show that this failure is not explained by the inadequate choice of transfer loss or by differences in pre-training recipes. Our findings refine the prevailing understanding of attention in ViT representations: attention is sufficient \textit{only} when the student architecture matches the teacher.
Huaiyuan Qin, Muli Yang, Gabriel James Goenawan +4
Institute for Infocomm Research (I2R), A*STAR, Singapore · 3Shanghai Jiao Tong University · 2Sichuan University
Vision transformers face significant computational overheads in high-resolution dense prediction due to the quadratic complexity of self-attention. Linear attention offers efficiency but sacrifices local context modeling. We propose \textbf{HSMLA (Hierarchical Softmax Multi-scale Linear Attention)}, which combines ReLU-based linear attention for global context, selective softmax refinement for critical local features, and multi-scale token representations via depthwise convolutions. HSMLA achieves superior accuracy-efficiency trade-offs: up to 4.2× inference-time speedup across dense prediction tasks, 87.3 Dice with 3.2× speedup on CT organ segmentation, and 94.2 AUC with 4.1× speedup on pathology WSI.
Dong Liu, Yanxuan Yu, Renata Borovica-Gajic +2
University of California, Los Angeles, USA · Columbia University, USA · The University of Melbourne, Australia +1