Increasing depth of transformer models improves recognition, but it comes at a substantial cost. Each additional layer requires more parameters, which makes the process computationally inefficient. We ask whether additional processing can focus on integrating features already computed. Looped Audio Spectrogram Transformer (LAST) first processes all tokens, then reuses the same blocks to refine only the class token over fixed audio features, thereby making later passes inexpensive. On AudioSet, ten-pass LAST achieves 0.345 mean average precision, exceeding a twelve-layer sequential transformer by 2.1% relative with 49.4% fewer parameters, 42% fewer multiply-accumulate operations, and 9.8% higher measured throughput. Across separately trained models, increasing the pass count from two to ten improves accuracy while adding only 1.2% computation. Further evaluations show improved robustness to temporal masking and various other auditory augmentations, with better generalization on classification tasks with music, environmental, and event sounds.
Figures & tables
Figure 1: LAST outperformed a deeper transformer at lower cost: AudioSet Computation per clip (x-axis) v/s AudioSet mAP (y-axis) Sequential AST uses D∈{3,6,9,12} layers; the recurrent models use D=6 shared blocks and R∈{2,4,6,8,10} passes.
Figure 2: Sequential AST uses distinct blocks once. Full-token recurrence reuses the same blocks and updates all tokens on every pass. LAST uses one full pass, then updates only the class token using fixed audio features. Recurrent models use six shared blocks.
Figure 3: LAST degraded more gracefully under temporal masking: AudioSet mAP under temporal masking. Masking zeros the corresponding normalized spectrogram frames. Models: sequential D=12 and recurrent D=6 , R=10 .
Figure 4: LAST improved performance across the board: AudioSet mAP under waveform corruption. Each dot is the mean over settings within one family; each dashed line is the mean across the twelve families. Models: sequential D=12 and LAST D=6 , R=10 .
Sequential
Looped,
LAST
AST
full-token
Depth
12
6
6
Recurrences
–
2
10
Parameters (M)
86.1
43.5
43.5
MACs (G)
48.4
48.4
28.2
Throughput (clips/s)
571
572
628
Table 1: LAST’s features transferred better on most tasks, with half the parameters: Frozen-feature classification accuracy (%) using final-epoch encoders and the protocols in Section 3.1 .