cs.CVAug 7, 2026

HSMLA: Hierarchical Softmax Multi-scale Linear Attention for Efficient Vision Transformers

Authors: Dong LiuYanxuan YuRenata Borovica-GajicTong GengYing Nian Wu

Organizations: University of California, Los Angeles, USA · Columbia University, USA · The University of Melbourne, Australia · Rice University, USA

Abstract

Vision transformers face significant computational overheads in high-resolution dense prediction due to the quadratic complexity of self-attention. Linear attention offers efficiency but sacrifices local context modeling. We propose \textbf{HSMLA (Hierarchical Softmax Multi-scale Linear Attention)}, which combines ReLU-based linear attention for global context, selective softmax refinement for critical local features, and multi-scale token representations via depthwise convolutions. HSMLA achieves superior accuracy-efficiency trade-offs: up to 4.2×4.2\times inference-time speedup across dense prediction tasks, 87.387.3% Dice with 3.2×3.2\times speedup on CT organ segmentation, and 94.294.2% AUC with 4.1×4.1\times speedup on pathology WSI.

Explore similar work

CardsList