Organizations: School of Electrical and Electronic Engineering, Nanyang Technological University, Singapore · Singapore Management University · Tencent Hy Frontier Lab, Singapore · Department of Computer Science and Technology, Beijing Institute of Technology, China · University of Surrey, UK
Spatial audio large language models (LLMs) enable embodied agents, wearable assistants, and immersive systems to recognize sound events, localize sources, and reason about their spatial relationships. However, existing spatial audio LLMs often rely on early fusion of acoustic and spatial features and source-agnostic token representations. These designs make it difficult to preserve the correspondence between individual sound events and their spatial attributes, particularly in multi-source scenes. To address this limitation, we propose SAIL, a Spatial Audio Intelligence framework with LLMs that preserves acoustic-spatial structure and source-level correspondence from audio encoding to LLM alignment. SAIL introduces a Disentangled Spatial Audio Transformer that represents Mel-spectrogram and interaural phase difference features as separate acoustic and spatial streams. Source-discriminative task queries further learn event, direction, and distance information for each source. A Dual-Stream Q-Former then aligns the two streams with the LLM using acoustic and spatial queries organized by source slots. Compared with the early-fusion baseline, SAIL achieves consistent improvements in dual-source sound event detection, direction and distance estimation, and spatial reasoning. These results demonstrate the importance of structured, source-discriminative audio representations for multi-source spatial understanding and reasoning.
Figures & tables
Fig. 2: Illustration of the proposed SAIL architecture. The DSAT encoder disentangles binaural audio into Mel-based acoustic tokens and IPD-based spatial tokens, while the Dual-Stream Q-Former uses dual-stream queries to extract spatial-audio tokens for each sound source. The resulting audio tokens are combined with text tokens and fed into the LoRA-adapted LLM for spatial audio-language reasoning.
Setting
Model
Input
mAP ( ↑ )
ER 20∘ ( ↓ )
MAE ( ↓ )
DER ( ↓ )
Single-source
AudioMAE [ 37 ]
Mel-spectrograms (mono)
47.18
-
-
-
SELDnet [ 23 ]
Mel-spectrograms, IPD
42.66
25.19
19.21
38.46
Spatial-AST [ 20 ]
Mel-spectrograms, IPD
49.86
23.97
18.03
32.96
DSAT (ours)
Mel-spectrograms, IPD
50.29
26.97
20.87
37.99
Dual-source: one source matched
Spatial-AST [ 20 ]
Mel-spectrograms, IPD
24.14
35.86
26.72
37.76
DSAT (ours)
Mel-spectrograms, IPD
31.35
27.29
22.27
30.41
TABLE I: Comparison of spatial audio encoders on single-source and dual-source SELD tasks. ”Dual-source: one source matched” scores each model only on its best-covered source, while ”Dual-source: both sources matched” scores predictions for both sources after permutation matching. Metrics include mean Average Precision (mAP ↑ ), Error Rate at 20∘ (ER 20∘↓ ), Mean Angular Error (MAE ↓ ), and Distance Error Rate (DER ↓ ).
Perception (Type ABCD)
Reasoning (Type E, Binary Acc ↑ )
Open-ended Reasoning ‡
Model
Input
Detection (mAP ↑ )
DoA (Acc ↑ )
DP (DER ↓ )
Direction
Distance
Avg.
Source (mAP ↑ )
Direction (Acc ↑ )
Inter-source DP (DER ↓ )
Random
–
0.65 ∣ 0.64
12.69 ∣ 12.82
65.53 ∣ 77.78
49.99
49.66
49.83
0.77
50.28
63.79
BAT [ 20 ]
B + P
24.51 ∣ 8.05
72.71 ∣ 35.12
34.13 ∣ 52.85
69.88
80.35
75.12
10.27
50.68
50.53
SAIL (ours)
P
0.61 ∣ 0.60
14.34 ∣ 13.62
59.26 ∣ 71.63
48.06
53.41
50.74
0.79
46.38
89.14
B + P
22.95 ∣ 10.02
76.81 ∣ 46.65
34.66 ∣ 47.25
76.13
86.74
81.44
12.69
62.78
50.54
TABLE II: Comparison of SAIL with BAT on SpatialSoundQA. Perception is evaluated on sound detection, DoA estimation, and distance prediction (DP). Values before and after “ ∣ ” denote single-source (Types A and B) and dual-source (Types C and D) results, respectively, with single-source results shaded in grey. Type-E reasoning includes binary Yes/No questions evaluated by Binary Accuracy (Acc) and additional open-ended dual-source questions on direction-conditioned source identification (Source, mAP), relative source direction (Direction, Acc), and inter-source distance (DP, DER).
Perception (Type ABCD)
Reasoning (Type E, Binary Acc ↑ )
Open-ended Reasoning
LLM Backbone
#Params
Detection (mAP ↑ )
DoA (Acc ↑ )
DP (DER ↓ )
Direction
Distance
Avg.
Source (mAP ↑ )
Direction (Acc ↑ )
Inter-source DP (DER ↓ )
Llama-2-7b
7B
22.95 ∣ 10.02
76.81 ∣ 46.65
34.66 ∣ 47.25
76.13
86.74
81.44
12.69
62.78
50.54
Llama-3.1-8B
8B
23.11 ∣ 9.01
76.71 ∣ 45.67
33.54 ∣ 47.53
74.24
82.48
78.36
11.78
60.30
47.71
TABLE III: Ablation on the LLM backbone of SAIL. All variants share the same DSAT encoder, Q-Former, and three-stage training curriculum, differing only in the underlying LLM. #Params denotes the total parameter count of the LLM backbone. Metrics and column conventions follow Table II : values before and after “ ∣ ” correspond to single-source and dual-source questions, respectively, with single-source entries shaded in grey.
Appendix figures & tables1 asset
Supplementary material from the paper’s appendix.
Appendix
DSAT Encoder Pre-training
SAIL Instruction Tuning
Configuration
Single-source Stage I
Single-source Stage II
Mixed-source
Stage I
Stage II
Stage III
Training objective
Detection
Detection, distance, DoA
Detection, distance, DoA
Single-source perception
Dual-source perception
Spatial reasoning
Trainable modules
DSAT
DSAT
DSAT
Dual-Stream Q-Former and LoRA
Training loss
Detection BCE
Multi-task
Multi-task
Language-modeling cross-entropy
λcls:λdist:λdoa
1000:0:0
400:4:2
400:8:2
Not applicable
Initialization
AudioMAE
Previous stage
Single-source DSAT
Random initialization
SAIL Stage I
SAIL Stage II
Appendix
TABLE IV: Training configurations for the DSAT encoder and SAIL.
Recent multimodal large language models mainly process audio as monaural signals, thereby discarding the spatial cues contained in spatial audio for sound localization, spatial relation reasoning, and spatial scene understanding. We propose Spatial-Omni, a lightweight method that implements SO-Encoder to inject First-Order Ambisonics (FOA) spatial audio into existing Omni LLMs as an independent modality, without modifying their original audio encoders. SO-Encoder provides spatial tokens with limited additional context cost and improves spatial audio understanding through efficient staged training. To support training and evaluation, we construct SO-Dataset, SO-QA, and SO-Bench from open-source data, real recordings, and simulations, containing 400K FOA spatial audio clips and 2.1M spatial question answering pairs. SO-Bench covers 16 spatial audio understanding subtasks, including basic detection and location estimation, spatial relation understanding, and complex spatial reasoning. Experiments show that Spatial-Omni outperforms existing open-source Large Audio-Language Models (LALMs) and Omni LLM models on spatial audio understanding tasks while retaining a reasonable level of general audio understanding. Code and data are available at https://github.com/dieKarotte/Spatial-Omni.
Sound events are entities with semantic identities, locations, and trajectories, but current audio-language models usually reason about clips as global event content. Conversely, sound event localization models track source directions over time but offer limited semantic coverage for language reasoning. To address this gap, we introduce ST-AudioQA, a spatio-temporal audio QA dataset and benchmark built from first-order ambisonic (FOA) renderings of static and moving sound sources. Each scene provides source identity, activity, direction, distance, and motion metadata, enabling dense trajectory supervision and questions about what is sounding, where it is, how it moves, and how sources relate. We further propose ST-Audio Encoder, a time-resolved FOA audio encoder that learns event semantics together with source trajectories, and ST-AudioLM, which connects the audio tokens from the encoder to an LLM for spatio-temporal audio QA. Experiments show that this representation improves the semantic-localization tradeoff and yields stronger reasoning performance than static spatial and localization-oriented baselines.
Embodied, ego-centric intelligence fundamentally requires the ability to comprehend spatial audio within complex environments. While large audio-language models excel at mono-channel reasoning, they lack spatial awareness, discarding critical spatial cues that enable sound localization and that can improve the disentanglement of overlapping sound sources. To address this, we present SEA-LM, a Spatial Audio Understanding model. First, we introduce FOACODER, a layout-flexible spatial audio encoder trained on source localization and ego-centric voice activity detection objectives to encode First Order Ambisonics derived from variable-count, variable-position smart-glasses arrays via beamforming. We train a Multimodal Large Language Model (MLLM) to understand these spatial audio embeddings through a two-stage curriculum spanning six tasks, including sound localization and spatially selective transcription in settings with multiple speakers and overlapping sounds. To prevent the transcription outputs from dominating the next token prediction loss and overwhelming the direction predictions, we introduce a Spatio-temporal Weighted Cross-Entropy Loss. On our evaluation set, SEA-LM achieves lower azimuth and elevation MAE, higher temporal IoU, lower external-source hallucination and missing-source rates, and lower WER on most transcription tasks than compared baselines, while remaining robust across 1,211 smart-glasses array configurations with 4 to 9 microphones.
Sonal Kumar, Sinan Hersek, Artem Dementyev +6
University of Maryland, College Park · Google · Google DeepMind