cs.SDOct 1, 2026

LAST: Looped Audio Spectrogram Transformer

Authors: Haider Al-Tahan, Sean O'Brien, Anastasia Razdaibiedina, N. Apurva Ratan Murty

Organizations: Georgia Institute of Technology · University of California San Diego · Google DeepMind

Abstract

Increasing depth of transformer models improves recognition, but it comes at a substantial cost. Each additional layer requires more parameters, which makes the process computationally inefficient. We ask whether additional processing can focus on integrating features already computed. Looped Audio Spectrogram Transformer (LAST) first processes all tokens, then reuses the same blocks to refine only the class token over fixed audio features, thereby making later passes inexpensive. On AudioSet, ten-pass LAST achieves 0.345 mean average precision, exceeding a twelve-layer sequential transformer by 2.1% relative with 49.4% fewer parameters, 42% fewer multiply-accumulate operations, and 9.8% higher measured throughput. Across separately trained models, increasing the pass count from two to ten improves accuracy while adding only 1.2% computation. Further evaluations show improved robustness to temporal masking and various other auditory augmentations, with better generalization on classification tasks with music, environmental, and event sounds.

Figures & tables

Explore similar work

CardsList
  1. Test-Time Compute Scaling for ASR with Depth-Conditioned Looped Transformers

    Jun 3, 2026Yacouba Kaloga, Shashi Kumar, Shakeel A. Sheikh +3Automatic Speech RecognitionTransformer Architectures

  2. Audio Token Attention Is Predictable Before the Language Model Runs

    Sep 30, 2026Kyoungjun Park, Yunzhe Li, Lili QiuLarge Audio Language ModelsAudio Understanding

  3. Inside the Latent Flow: Causal Deciphering of Attention Dynamics in Audio Separation Foundation Models

    Jun 8, 2026Yuxuan Chen, Haoyuan Yu, Peize HeSpeech SeparationAudio Understanding