cs.CVOct 7, 2026

On the Necessity of Attention-FFN Split in Vision Transformers

Authors: Junhyeok Kim, Jinyeong Kim, Jae Wan Park, Seong Jae Hwang

Organizations: Department of Artificial Intelligence, Yonsei University, Seoul, Republic of Korea.

Abstract

The standard Transformer architecture relies on a rigid pattern that alternates Attention and Feed-Forward Network (FFN) layers. Despite its widespread adoption, the inductive bias imposed by this strict separation has not been systematically examined. In this work, we investigate the necessity of the Attention-FFN dichotomy in Vision Transformers (ViTs). To facilitate this analysis, we introduce the AttenFeed module, a unified component that integrates the functional properties of both Attention and FFN. Based on this module, we devise the unified Vision Transformer (uViT), which replaces the conventional alternating Attention-FFN structure with a sequence of AttenFeed modules. We then use uViT as a control group that relaxes the Attention-FFN dichotomy of the standard ViT and systematically compare the two models across multiple datasets and model scales. Our experiments reveal that the Attention-FFN dichotomy can hinder performance at smaller model scales due to the rigid parameter allocation of ViTs. The AttenFeed module and uViT serve as new analytical tools for understanding the Attention-FFN structure and offer theoretical insights into the heuristically designed architecture of conventional ViTs.

Figures & tables

Appendix figures & tables4 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Attention Transfer Is Not Universally Effective for Vision Transformers

    May 8, 2026Huaiyuan Qin, Muli Yang, Gabriel James Goenawan +4Self-AttentionVision Transformer

  2. Foveation-Guided Dynamic Token Selection for Robust and Efficient Vision Transformers

    Jul 10, 2026Ibrahim Batuhan Akkaya, Kishaan Jeeveswaran, Bahram Zonooz +1Image Corruption RobustnessEfficient ViTs

  3. Do Vision Transformers Need All-to-All Attention? Global Communication Through Elastic Learned Cores

    May 12, 2026Alan Z. Song, Yinjie Chen, Mu Nan +3Visual Representation LearningVision Transformer