cs.CVApr 19, 2026

Attention Is not Everything: Efficient Alternatives for Vision

Authors: Nur Mohammad KaziIbteshum KhaledMd. Luthful Hasan GalibAli Faruk ShihabMd. Rakibul Islam

Organizations: Ahsanullah University of Science and Technology Dhaka, Bangladesh

Abstract

Recently computer vision has seen advancements mainly thanks to Transformer-based models. However many non-Transformer methods are still doing well being a direct competition of Transformer-based models. This review tries to present a comprehensive taxonomy of such methods and organize these methods into categories like convolution-based models, MLP-based models, state-space-based and more. These methods are looked at in terms of how efficient they are, how well they scale, how easy they are to understand and how robust they are. A total of 40 papers were chosen for this study. The goal is to give a view of non-Transformer methods and find out what challenges and opportunities exist for future computer vision research.

Explore similar work

CardsList
  1. A Smaller Transformer in Your Transformer

    Sep 17, 2026Dhananjay Tomar, Marius Aasan, Andreas Kleppe +1Vision TransformerTransformer Architectures