Copy the Same, Distill the Difference: Initializing Linear Vision Transformers
Authors: Huaiyuan Qin, Muli Yang, Gabriel James Goenawan, Shiqi Huang, Min Kass Chong, Wahyu Wiratama, Peng Hu, Chen Gong, +4 more
Organizations: Institute of Advanced Intelligence and Computing (IAIC), A*STAR, Singapore · Nanyang Technological University · ST Engineering Geo-Insights, Singapore · Sichuan University · Shanghai Jiao Tong University · University of Science and Technology of China
Linear Vision Transformers (ViTs) are designed to replace the attention in Softmax ViTs with the linear-complexity attention operator for more efficient token routing, but they require from-scratch pre-training and typically underperform the original Softmax version. How to initialize linear ViTs both efficiently and effectively still remains unclear. In this work, we explicitly ask: given that most foundation ViTs are built on the mainstream Softmax attention, can linear ViTs benefit from their pre-trained weights? Recent works on Attention Transfer show that attention is the effective transferable component between Softmax ViTs, suggesting attention alone suffices for such reuse. However, we find the opposite for Softmax-to-linear transfer. The attention weights are operator-specific: copying them barely helps, and is sometimes even worse than random initialization. Instead, the attention's token routing behavior can be recovered through distillation with a proper loss design, letting linear ViTs reduce the gap and even match Softmax ones. In contrast, the MLP weights, which carry the learned representation, are operator-agnostic: they can be transferred by simple direct copying, which already carries most of the benefit of the pre-trained weights. Thus, copying MLPs can serve as an effective foundation for Softmax-to-linear transfer: paired with the distilled attention, linear ViTs eventually close the remaining gap and even surpass Softmax ones. These findings hold consistently across various linear ViT variants, different model sizes, and diverse datasets. We hope this study deepens the understanding of reusing pre-trained weights across attention operators: copy what stays the same and distill what differs, to recover the benefit across the Softmax-to-linear boundary.
Figures & tables
Figure 1: What survives the Softmax-to-linear boundary? We ask whether linear ViTs can benefit from pre-trained Softmax weights for initialization. We find that the attention weights are operator-specific that copying them barely helps. Their token routing behavior can instead be recovered by distilling the attention outputs, while the operator-agnostic MLP weights, which carry the learned representation, can be simply transferred by direct copying. Paired together, linear ViTs can match and even surpass the Softmax ones, consistently across diverse settings.
ELU
ReLU
MHLA
TTT
Random
79.6
76.5
81.8
82.2
DeiT
79.8 +0.2
75.6 -0.9
81.7 -0.1
81.0 -1.2
DINO
80.2 +0.6
75.2 -1.3
81.6 -0.2
80.6 -1.6
MoCov3
79.9 +0.3
75.7 -0.8
81.7 -0.1
81.1 -1.1
iBOT
80.3 +0.7
75.3 -1.2
81.8 +0.0
80.7 -1.5
MAE
80.4 +0.8
76.0 -0.5
82.0 +0.2
81.4 -0.8
Table 1: Pre-trained weights robustness. We repeat the attention-only transfer on ViT-Base with different pre-trained Softmax weights, under the same ImageNet-1K evaluation. Gray row shows the shared Random baseline, and colored deltas indicate gains/drops relative to it. The failure remains consistent across all pre-trained sources: none provides a reliable advantage over random initialization.
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
Setting
Copied components
Class token
Pos. embedding
Random
–
random
random
Mimetic
–
random
random
Impulse
–
random
random
Copy
Attention
random
random
MLPCopy
MLPs
random
random
FullCopy
all Softmax weights
copied
copied
Appendix
Table A: Implementation details of transfer protocol. We list the weight-initialization details of each component in the Softmax-to-linear transfer experiments. Note that we follow the default implementations in MHLA and TTT to use sinusoidal positional embeddings and average pooling with no class token. All other settings keep the class token and learned positional embeddings. Mimetic and Impulse assume the standard QK⊤ attention parameterization and are therefore incompatible with TTT.
Dataset
Classes
Train
Test / Val
Epochs
Warm-up
ImageNet-1K ( Deng et al., 2009 )
1,000
1,281,167
50,000
300
50
CIFAR-10 ( Krizhevsky, 2009 )
10
50,000
10,000
300
50
CIFAR-100 ( Krizhevsky, 2009 )
100
50,000
10,000
300
50
STL-10 ( Coates et al., 2011 )
10
5,000
8,000
300
50
Food ( Bossard et al., 2014 )
101
75,750
25,250
300
50
Flowers ( Nilsback and Zisserman, 2008 )
102
1,020
6,149
600
100
Appendix
Table B: Details of datasets and training schedules.
Config
Transfer datasets
ImageNet-1K
Optimizer
AdamW
AdamW
Base Learning Rate
2e-3
2e-3
Weight Decay
0.05
0.05
Optimizer Momentum
β1=0.9 , β2=0.999
β1=0.9 , β2=0.999
Layer-wise LR Decay
–
–
Batch Size
512
4096
Appendix
Table C: Training recipes for transfer datasets and ImageNet-1K.
w/
IN-1K
COCO
ADE20K
Distill
Acc
AP b
AP 50b
AP 75b
AP m
AP 50m
AP 75m
mIoU
+MS
Softmax
81.8
42.9
65.7
46.8
39.4
62.6
42.0
46.1
47.1
Random
76.5
34.0
56.5
37.6
32.1
55.5
34.4
34.9
35.7
✓
79.4 +2.9
36.2 +2.2
59.0 +2.5
40.3 +2.7
35.8 +3.7
59.2 +3.7
37.9 +3.5
40.7 +5.8
41.8 +6.1
Copy
75.6
33.3
55.9
37.1
31.6
55.0
33.7
35.2
35.8
✓
79.9 +4.3
36.9 +3.6
59.3 +3.4
40.7 +3.6
36.0 +4.4
59.6 +4.6
38.2 +4.5
41.5 +6.3
42.1 +6.3
Appendix
Table D: ImageNet-1K and dense prediction results with ReLU at Base size. All settings adopt transferred weights from DeiT and are then fine-tuned with Mask R-CNN (1 × ) on COCO and UperNet (160k iters) on ADE20K. We report the Top-1 Accuracy (%) on ImageNet-1K, box and mask AP on COCO, and mIoU with single and multi-scale (+MS) on ADE20K, respectively. w/ Distill with ✓ indicates applying attention distillation during the Softmax-to-linear transfer, and colored deltas indicate gains relative to the same setting w/o distillation. Softmax row refers to the Softmax reference and the plain-ViT baselines of Chen et al. (2023) .
Weight Copying
Accuracy
Attn.
MLP
Patch emb.
Pos. emb.
Cls. token
w/o Distill
w/ Distill
FullCopy
✓
✓
✓
✓
✓
80.8 +4.3
82.0 +2.6
Variant 1
✓
✓
✓
✓
80.7 +4.2
81.8 +2.4
Variant 2
✓
✓
✓
80.6 +4.1
81.9 +2.5
Variant 3
✓
✓
80.7 +4.2
81.9 +2.5
MLPCopy
✓
80.2 +3.7
81.4 +2.0
Appendix
Table G: Ablation on copying the remaining components. Starting from FullCopy , we gradually remove the weight loading of the class token, positional embedding, and patch embedding, while all attention and MLP weights remain copied. We report the Top-1 Accuracy (%) on ImageNet-1K with ReLU at Base size, with colored deltas relative to the Random baseline.