cs.SDSep 28, 2026

SAIL: Spatial Audio Intelligence with Large Language Models via Disentangled Acoustic-Spatial Encoding and Dual-Stream Q-Former

Authors: Zhengding Luo, Jinyang Wu, Haozhe Ma, Yanghao Zhou, Woon-Seng Gan, Wenwu Wang

Organizations: School of Electrical and Electronic Engineering, Nanyang Technological University, Singapore · Singapore Management University · Tencent Hy Frontier Lab, Singapore · Department of Computer Science and Technology, Beijing Institute of Technology, China · University of Surrey, UK

Abstract

Spatial audio large language models (LLMs) enable embodied agents, wearable assistants, and immersive systems to recognize sound events, localize sources, and reason about their spatial relationships. However, existing spatial audio LLMs often rely on early fusion of acoustic and spatial features and source-agnostic token representations. These designs make it difficult to preserve the correspondence between individual sound events and their spatial attributes, particularly in multi-source scenes. To address this limitation, we propose SAIL, a Spatial Audio Intelligence framework with LLMs that preserves acoustic-spatial structure and source-level correspondence from audio encoding to LLM alignment. SAIL introduces a Disentangled Spatial Audio Transformer that represents Mel-spectrogram and interaural phase difference features as separate acoustic and spatial streams. Source-discriminative task queries further learn event, direction, and distance information for each source. A Dual-Stream Q-Former then aligns the two streams with the LLM using acoustic and spatial queries organized by source slots. Compared with the early-fusion baseline, SAIL achieves consistent improvements in dual-source sound event detection, direction and distance estimation, and spatial reasoning. These results demonstrate the importance of structured, source-discriminative audio representations for multi-source spatial understanding and reasoning.

Figures & tables

Appendix figures & tables1 asset

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Spatial-Omni: Spatial Audio Understanding Integration in Multimodal LLMs via FOA Encoding

    Jun 9, 2026Zhiyuan Zhu, Yixuan Chen, Yiwen Shao +13Spatial AudioOmni-Modal

  2. Spatio-Temporal Audio Language Modeling for Dynamic Sound Sources

    Jun 12, 2026Oh Hyun-Bin, Kazuki Shimada, Yuhta Takida +6Sound Source LocalizationLarge Audio Language Models

  3. SEA-LM: Egocentric Spatial Audio Understanding for Wearable Microphone Arrays

    Oct 4, 2026Sonal Kumar, Sinan Hersek, Artem Dementyev +6