Organizations: School of Electrical and Electronic Engineering, Nanyang Technological University, Singapore · Singapore Management University · Tencent Hy Frontier Lab, Singapore · Department of Computer Science and Technology, Beijing Institute of Technology, China · University of Surrey, UK
Spatial audio large language models (LLMs) enable embodied agents, wearable assistants, and immersive systems to recognize sound events, localize sources, and reason about their spatial relationships. However, existing spatial audio LLMs often rely on early fusion of acoustic and spatial features and source-agnostic token representations. These designs make it difficult to preserve the correspondence between individual sound events and their spatial attributes, particularly in multi-source scenes. To address this limitation, we propose SAIL, a Spatial Audio Intelligence framework with LLMs that preserves acoustic-spatial structure and source-level correspondence from audio encoding to LLM alignment. SAIL introduces a Disentangled Spatial Audio Transformer that represents Mel-spectrogram and interaural phase difference features as separate acoustic and spatial streams. Source-discriminative task queries further learn event, direction, and distance information for each source. A Dual-Stream Q-Former then aligns the two streams with the LLM using acoustic and spatial queries organized by source slots. Compared with the early-fusion baseline, SAIL achieves consistent improvements in dual-source sound event detection, direction and distance estimation, and spatial reasoning. These results demonstrate the importance of structured, source-discriminative audio representations for multi-source spatial understanding and reasoning.
Figures & tables
Fig. 2: Illustration of the proposed SAIL architecture. The DSAT encoder disentangles binaural audio into Mel-based acoustic tokens and IPD-based spatial tokens, while the Dual-Stream Q-Former uses dual-stream queries to extract spatial-audio tokens for each sound source. The resulting audio tokens are combined with text tokens and fed into the LoRA-adapted LLM for spatial audio-language reasoning.
Setting
Model
Input
mAP ( ↑ )
ER 20∘ ( ↓ )
MAE ( ↓ )
DER ( ↓ )
Single-source
AudioMAE [ 37 ]
Mel-spectrograms (mono)
47.18
-
-
-
SELDnet [ 23 ]
Mel-spectrograms, IPD
42.66
25.19
19.21
38.46
Spatial-AST [ 20 ]
Mel-spectrograms, IPD
49.86
23.97
18.03
32.96
DSAT (ours)
Mel-spectrograms, IPD
50.29
26.97
20.87
37.99
Dual-source: one source matched
Spatial-AST [ 20 ]
Mel-spectrograms, IPD
24.14
35.86
26.72
37.76
DSAT (ours)
Mel-spectrograms, IPD
31.35
27.29
22.27
30.41
TABLE I: Comparison of spatial audio encoders on single-source and dual-source SELD tasks. ”Dual-source: one source matched” scores each model only on its best-covered source, while ”Dual-source: both sources matched” scores predictions for both sources after permutation matching. Metrics include mean Average Precision (mAP ↑ ), Error Rate at 20∘ (ER 20∘↓ ), Mean Angular Error (MAE ↓ ), and Distance Error Rate (DER ↓ ).
Perception (Type ABCD)
Reasoning (Type E, Binary Acc ↑ )
Open-ended Reasoning ‡
Model
Input
Detection (mAP ↑ )
DoA (Acc ↑ )
DP (DER ↓ )
Direction
Distance
Avg.
Source (mAP ↑ )
Direction (Acc ↑ )
Inter-source DP (DER ↓ )
Random
–
0.65 ∣ 0.64
12.69 ∣ 12.82
65.53 ∣ 77.78
49.99
49.66
49.83
0.77
50.28
63.79
BAT [ 20 ]
B + P
24.51 ∣ 8.05
72.71 ∣ 35.12
34.13 ∣ 52.85
69.88
80.35
75.12
10.27
50.68
50.53
SAIL (ours)
P
0.61 ∣ 0.60
14.34 ∣ 13.62
59.26 ∣ 71.63
48.06
53.41
50.74
0.79
46.38
89.14
B + P
22.95 ∣ 10.02
76.81 ∣ 46.65
34.66 ∣ 47.25
76.13
86.74
81.44
12.69
62.78
50.54
TABLE II: Comparison of SAIL with BAT on SpatialSoundQA. Perception is evaluated on sound detection, DoA estimation, and distance prediction (DP). Values before and after “ ∣ ” denote single-source (Types A and B) and dual-source (Types C and D) results, respectively, with single-source results shaded in grey. Type-E reasoning includes binary Yes/No questions evaluated by Binary Accuracy (Acc) and additional open-ended dual-source questions on direction-conditioned source identification (Source, mAP), relative source direction (Direction, Acc), and inter-source distance (DP, DER).
Perception (Type ABCD)
Reasoning (Type E, Binary Acc ↑ )
Open-ended Reasoning
LLM Backbone
#Params
Detection (mAP ↑ )
DoA (Acc ↑ )
DP (DER ↓ )
Direction
Distance
Avg.
Source (mAP ↑ )
Direction (Acc ↑ )
Inter-source DP (DER ↓ )
Llama-2-7b
7B
22.95 ∣ 10.02
76.81 ∣ 46.65
34.66 ∣ 47.25
76.13
86.74
81.44
12.69
62.78
50.54
Llama-3.1-8B
8B
23.11 ∣ 9.01
76.71 ∣ 45.67
33.54 ∣ 47.53
74.24
82.48
78.36
11.78
60.30
47.71
TABLE III: Ablation on the LLM backbone of SAIL. All variants share the same DSAT encoder, Q-Former, and three-stage training curriculum, differing only in the underlying LLM. #Params denotes the total parameter count of the LLM backbone. Metrics and column conventions follow Table II : values before and after “ ∣ ” correspond to single-source and dual-source questions, respectively, with single-source entries shaded in grey.
Appendix figures & tables1 asset
Supplementary material from the paper’s appendix.
Appendix
DSAT Encoder Pre-training
SAIL Instruction Tuning
Configuration
Single-source Stage I
Single-source Stage II
Mixed-source
Stage I
Stage II
Stage III
Training objective
Detection
Detection, distance, DoA
Detection, distance, DoA
Single-source perception
Dual-source perception
Spatial reasoning
Trainable modules
DSAT
DSAT
DSAT
Dual-Stream Q-Former and LoRA
Training loss
Detection BCE
Multi-task
Multi-task
Language-modeling cross-entropy
λcls:λdist:λdoa
1000:0:0
400:4:2
400:8:2
Not applicable
Initialization
AudioMAE
Previous stage
Single-source DSAT
Random initialization
SAIL Stage I
SAIL Stage II
Appendix
TABLE IV: Training configurations for the DSAT encoder and SAIL.