cs.SDSep 28, 2026

On Temporal Binding in Large Audio Language Models

Authors: Paul Primus, Gerhard Widmer

Organizations: Institute of Computational Perception, Johannes Kepler University Linz, Austria · LIT Artificial Intelligence Lab, Johannes Kepler University Linz, Austria

Abstract

Reasoning about temporal structure of audio recordings requires Large Audio Language Models (LALMs) to associate sound events with their temporal position. Understanding the underlying mechanisms is a first step toward diagnosing failures and identifying model components that may need improvement. Using mechanistic interpretability, we investigate how temporal information is represented and bound to sound events in three open-source LALMs. We find that across all three, event-specific location becomes concentrated in event name representations at intermediate modality integration layers. These representations encode coarse event position along a low-dimensional, curved relative time trajectory. Steering event name representations along this trajectory systematically shifts before/after beliefs, providing evidence that these representations contribute to coarse temporal reasoning. In contrast, the same interventions do not reliably shift predicted onset timestamps, suggesting that coarse temporal reasoning and precise metric event localization rely on distinct mechanisms.

Explore similar work

CardsList
  1. A Closer Look at Failure Modes in Temporal Understanding of Large Audio-Language Models

    Jun 16, 2026Apoorva Kulkarni, Kaousheik Jayakumar, Sreyan Ghosh +3Large Audio Language ModelsAudio Understanding

  2. Augmenting Large Audio-Language Models with Frame-Level Grounding for Fine-Grained Temporal Perception

    Sep 14, 2026Yanfeng Shi, Yan Song, Junhui Li +4Large Audio Language ModelsPerception

  3. TEMPO: Temporally-grounded Multi-task Post-training for Large Audio-Language Models

    Aug 30, 2026Apoorva Kulkarni, Kaousheik Jayakumar, Sreyan Ghosh +3Large Audio Language ModelsNeural Audio