Equivariant networks embed geometric symmetries as structural priors through weight sharing, achieving remarkable parameter efficiency across vision tasks. However, this parameter efficiency does not translate into compute efficiency: most existing implementations unroll the structured weights into dense matrices and dispatch them to generic dense kernels, so an equivariant layer costs no fewer MACs than its non-equivariant counterpart. In this paper, we observe that the equivariant linear (EQ-Linear) layer---the most fundamental and frequently used module in modern equivariant architectures---is essentially a circular convolution along the group dimension composed with a linear transform along the channel dimension. Building on this observation, we propose Flash EQ-Linear, an exact acceleration algorithm that reduces the cost to
2(T−1)/T2 of the original dense formulation (
T is the equivariant group size) by combining the Fourier convolution theorem along the group dimension with the conjugate symmetry of the real DFT. To translate these computational savings into wall-clock speedups, we further develop dedicated CUDA kernels for the
p4 group. At the operator level, Flash EQ-Linear achieves up to
2.1× forward speedup over PyTorch's highly optimized F.linear; at the network level, Flash EQ-ViT achieves up to
1.7× end-to-end speedup over both equivariant and non-equivariant baselines. As an operator-level acceleration algorithm, Flash EQ-Linear provides plug-and-play acceleration for diverse pretrained equivariant models, including EQ-ViT, EQ-Swin, EQ-VMamba, and EQ-INR, without retraining or architectural changes. Code is available at https://github.com/zhongchenzhao/FlashEQLinear.