Embodied, ego-centric intelligence fundamentally requires the ability to comprehend spatial audio within complex environments. While large audio-language models excel at mono-channel reasoning, they lack spatial awareness, discarding critical spatial cues that enable sound localization and that can improve the disentanglement of overlapping sound sources. To address this, we present SEA-LM, a Spatial Audio Understanding model. First, we introduce FOACODER, a layout-flexible spatial audio encoder trained on source localization and ego-centric voice activity detection objectives to encode First Order Ambisonics derived from variable-count, variable-position smart-glasses arrays via beamforming. We train a Multimodal Large Language Model (MLLM) to understand these spatial audio embeddings through a two-stage curriculum spanning six tasks, including sound localization and spatially selective transcription in settings with multiple speakers and overlapping sounds. To prevent the transcription outputs from dominating the next token prediction loss and overwhelming the direction predictions, we introduce a Spatio-temporal Weighted Cross-Entropy Loss. On our evaluation set, SEA-LM achieves lower azimuth and elevation MAE, higher temporal IoU, lower external-source hallucination and missing-source rates, and lower WER on most transcription tasks than compared baselines, while remaining robust across 1,211 smart-glasses array configurations with 4 to 9 microphones.
Figures & tables
Figure 1: Overview of SEA-LM and FOACoder architectures. (Left) FOACoder : a causal transformer encoder ingests a multi-channel complex spectrogram and produces encoded spatial tokens. (Right) SEA-LM : a dual-pathway MLLM that combines frozen FOACoder spatial features with a frozen semantic audio encoder, injecting both into an LLM for joint spatial and semantic reasoning over FOA audio and text queries. Tensor dimensions use B for batch size, F for STFT frequency bins, T for time frames, C for input audio channels, L for the number of temporal patches or tokens, and D for the embedding dimension; P denotes the temporal patch length and N the number of transformer blocks.
Metric
Definition
T1
T2
T3
T4
T5
T6
Format adherence
Percentage of outputs that parse and satisfy the required schema.
✓
✓
✓
✓
✓
✓
External-source hallucination and missing-source rates
Percentage of validly parsed examples with at least one unmatched predicted or ground-truth external source.
✓
✓
✓
✓
✗
✓
Source-type accuracy
Exact agreement between the sound descriptors of matched predicted and ground-truth external sources after text normalization and mapping synonyms to canonical labels.
✓
✗
✓
✗
✗
✗
Azimuth, elevation, and distance MAE; percentage distance error
Spatial localization errors over matched predicted and ground-truth external sources.
✓
✓
✓
✓
✗
✓
Location- and distance-bin accuracy
Exact agreement between the categorical spatial bins of matched predicted and ground-truth external sources.
✓
✓
✓
✓
✗
✓
Start- and end-time MAE; temporal IoU
Temporal boundary error and interval overlap over matched predicted and ground-truth sources.
✓
✓
✓
✓
✓
✓
Table 1: Evaluation metrics by task: T1, Holistic Sound Localization; T2, Selective Sound Localization; T3, Holistic Captioning; T4, Role-Aware Transcription; T5, Targeted Wearer Speech Transcription; and T6, Targeted External Speech Transcription. Accuracy, precision, recall, and F1 are reported for both wearer and query detection where applicable.
(a) Spatial and event understanding (Tasks 1–2).
Task 1: Holistic Sound Localization
Task 2: Selective Sound Localization
Model
Azim. ↓
Elev. ↓
Dist. ↓
Type acc. ↑
tIoU ↑
Azim. ↓
Elev. ↓
Query F1 † ↑
( ∘ )
( ∘ )
(m)
(%)
( ∘ )
( ∘ )
(%)
Gemma-4-E4B-it (vanilla)
67.9
27.7
1.88
0.9
0.37
48.0
26.4
10.9
Gemma-4-E4B-it (fine-tuned)
87.5
27.7
1.28
62.2
0.91
88.2
28.3
78.0
SELDNet + Gemma-4-E4B-it
80.6
27.4
1.26
55.6
0.61
80.2
27.4
76.2
Table 2: Task-wise results on the evaluation set. Best values are bold ; second-best values are underlined . Note: Azim., Elev., Dist., Type acc., and tIoU denote azimuth MAE, elevation MAE, distance MAE, source-type accuracy, and temporal intersection over union, respectively. Halluc. and Missing denote the external-source hallucination and missing-source rates. † Query F1 is the F1 detection score for the queried source type in Task 2 or the queried direction in Task 6.
Appendix figures & tables18 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 2: Spherical coordinate system showing distance r , elevation angle θ (measured from the xy -plane), and azimuth angle ϕ .
Figure 3: Ambisonic Signal Matching Simulation Results. (a) A 4-microphone array configuration on the Design 1 smart glasses (the microphone indices are visualized in Figure 4(a) ), with microphones located at the front and rear of both the left and right temples. Because these microphones are strictly coplanar, the elevation (Z) channel cannot be accurately reconstructed, resulting in significant attenuation across the RMS power spectrum and failure to meet our Normalized Mean Squared Error (NMSE) selection criteria. (b) A 5-microphone configuration that adds a single microphone on the bottom-right rim of the glasses. This addition introduces necessary Z-axis displacement, allowing the array to reconstruct the elevation channel sufficiently and pass the selection criteria. In both configurations, the inherent low-pass filtering effect of ASM beamforming is evident across most reconstructed channels.
(a) DCASE Metrics
Model
Parameter Count
Angular Error ( ∘ ) ↓
Error Rate ↓
F-score ↑
Localization Recall ↑
SELD Score ↓
SELDNet
0.74M
44.01
1.20
12.62
32.29
0.75
FOACoder 1s context
46.03M
15.84
0.41
63.43
75.87
0.28
FOACoder 30s context
46.18M
15.06
0.34
70.46
86.50
0.21
Appendix
Table 3: Encoder evaluation results. Best values in each column are indicated in bold .
Figure 4: Smart-glasses microphone array designs and simulation geometry. (a) Microphone placement on the Design 1 glasses array. (b) The Design 1 glasses fitted on a HEAD acoustics HMS II head-and-torso (HAT) model for array transfer function simulation. (c–d) Complementary views of the Design 2 glasses array topology. In both topology designs, numbered red arrows identify the microphone positions referenced in the array-selection analysis.
Scene Content
Holistic Sound Localization (Task 1)
Selective Sound Localization (Task 2)
Holistic Captioning (Task 3)
Role-Aware Transcription (Task 4)
Targeted Wearer Speech Transcription (Task 5)
Targeted External Speech Transcription (Task 6)
External Non-Speech Sounds
✓
✓
✓
✗
✗
✗
Wearer Speech
✗
✗
✓
✓
✓
✗
Bystander Speech
✓
✗
✓
✓
✗
✓
Appendix
Table 4: Task validity matrix based on scene content. A checkmark (✓) indicates that the presence of the specified sound source satisfies the requirements for the corresponding task to be valid.
Metric group
Holistic Sound Localization (Task 1)
Selective Sound Localization (Task 2)
Holistic Captioning (Task 3)
Role-Aware Transcription (Task 4)
Targeted Wearer Speech Transcription (Task 5)
Targeted External Speech Transcription (Task 6)
Format Adherence
✓
✓
✓
✓
✓
✓
Hallucinated/Missing External Sources
✓
✓
✓
✓
✗
✓
Source-Type Accuracy
✓
✗
✓
✗
✗
✗
Spatial Errors and Bin Accuracies
✓
✓
✓
✓
✗
✓
Temporal Errors and IoU
✓
✓
✓
✓
✓
✓
Speech Transcription (WER)
✗
✗
✓
✓
✓
✓
Appendix
Table 5: Applicability of metric groups to the six evaluation tasks. “External” refers to sources other than the wearer.
Figure 5: Results segmented by total source count. (a) Holistic Localization Error: Task 1 azimuth and elevation MAEs across total-source-count cohorts. (b) Holistic Localization Quality: Task 1 location-bin accuracy and temporal IoU across the same cohorts. (c) Transcription by Task: WER for Holistic Captioning, Role-Aware Transcription, Targeted Wearer Transcription, and Targeted External Transcription. (d) Query Detection: Query F1 for Selective Localization and Targeted External Transcription.
Figure 6: Results segmented by speaker count. Speaker count includes wearer- and bystander-speech sources in the complete scene. (a) Multi-Source Transcription: WER for Holistic Captioning and Role-Aware Transcription across speaker-count cohorts. (b) Targeted Transcription: WER for Targeted Wearer Transcription and Targeted External Transcription. (c) Temporal Grounding: Temporal IoU for the four transcription tasks. (d) Role and Query Detection: Wearer F1 for Holistic Captioning, Role-Aware Transcription, and Targeted Wearer Transcription, together with query F1 for Targeted External Transcription. The three wearer-F1 traces overlap almost completely near 100% , so they partially obscure one another.
Figure 7: Results segmented by source-overlap percentage. Overlap includes all speech and non-speech sources and is normalized by the union of active source intervals as defined in Equation 16 . (a) Holistic Localization Error: Task 1 azimuth and elevation MAEs across overlap cohorts. (b) Holistic Localization Quality: Task 1 location-bin accuracy and temporal IoU. (c) Transcription by Task: WER for Holistic Captioning, Role-Aware Transcription, Targeted Wearer Transcription, and Targeted External Transcription. (d) Query Detection: Query F1 for Selective Localization and Targeted External Transcription.
Figure 8: Results segmented by array microphone count. A single SEA-LM checkpoint is evaluated without array-specific retraining across scenes synthesized with 4–9 microphones. (a) Holistic Localization Error: Task 1 azimuth and elevation MAEs across microphone counts. (b) Holistic Localization Quality: Task 1 location-bin accuracy and temporal IoU. (c) Transcription by Task: WER for Holistic Captioning, Role-Aware Transcription, Targeted Wearer Transcription, and Targeted External Transcription. (d) Query Detection: Query F1 for Selective Localization and Targeted External Transcription.
Location Bin
Elevation Range ( θ )
Azimuth Range ( ϕ )
above
(80∘,90∘]
[0∘,360∘)
below
[−90∘,−80∘)
[0∘,360∘)
front
[−80∘,80∘]
(337.5∘,360∘)∪[0∘,22.5∘]
left
[−80∘,80∘]
(22.5∘,157.5∘]
rear
[−80∘,80∘]
(157.5∘,202.5∘]
right
[−80∘,80∘]
(202.5∘,337.5∘]
Appendix
Table 6: Categorical direction bins and their corresponding elevation ( θ ) and azimuth ( ϕ ) ranges in degrees.
Task 1: Holistic Sound Localization
Task 2: Selective Sound Localization
Curriculum
Azim. ↓
Elev. ↓
Dist. ↓
Loc. bin ↑
Type acc. ↑
tIoU ↑
Azim. ↓
Elev. ↓
Query F1 ↑
( ∘ )
( ∘ )
(m)
(%)
(%)
( ∘ )
( ∘ )
(%)
Stage 1 → Stage 2a → Stage 2
15.5
13.1
1.20
85.5
61.4
0.84
19.7
15.2
83.9
Stage 1 → Stage 2 (ours)
14.9
12.8
1.07
84.1
63.7
0.92
17.5
14.8
94.9
Appendix
Table 7: Effect of inserting focused Stage 2a between Stage 1 and Stage 2 on the held-out synthetic evaluation set. Arrows indicate the preferred direction for each metric.
Model
N
Format ↑
Halluc. ↓
Missing ↓
Type acc. ↑
(%)
(%)
(%)
(%)
Gemma-4-E4B-it (vanilla)
26,985
100.0
4.0
94.8
0.9
Gemma-4-E4B-it (fine-tuned)
26,985
100.0
6.8
10.6
62.2
SELDNet + Gemma-4-E4B-it
26,985
100.0
50.3
22.6
55.6
SEA-LM (ours)
26,985
100.0
3.6
5.8
63.7
Appendix
Table 8: Complete Task 1 (Holistic Sound Localization) results on the evaluation set. N = total evaluated examples; Format = format adherence; Halluc. = external-source hallucination rate; Missing = missing external source rate; Type acc. = source-type accuracy; Azim. = azimuth MAE; Elev. = elevation MAE; Dist. = distance MAE; Dist. err. = percentage distance error; Loc. bin = location-bin accuracy; Dist. bin = distance-bin accuracy; Start = start-time MAE; End = end-time MAE; tIoU = temporal intersection over union.
Model
N
Format ↑
Halluc. ↓
Missing ↓
(%)
(%)
(%)
Gemma-4-E4B-it (vanilla)
17,247
100.0
26.1
64.7
Gemma-4-E4B-it (fine-tuned)
17,247
100.0
34.0
4.0
SELDNet + Gemma-4-E4B-it
17,247
100.0
21.1
13.8
SEA-LM (ours)
17,247
100.0
4.5
2.9
Appendix
Table 9: Complete Task 2 (Selective Sound Localization) results on the evaluation set. N = total evaluated examples; Format = format adherence; Halluc. = external-source hallucination rate; Missing = missing external source rate; Azim. = azimuth MAE; Elev. = elevation MAE; Dist. = distance MAE; Dist. err. = percentage distance error; Loc. bin = location-bin accuracy; Dist. bin = distance-bin accuracy; Start = start-time MAE; End = end-time MAE; tIoU = temporal intersection over union; Detection acc. = source-level query-detection accuracy; Precision = precision; Recall = recall; F1 = F1 score.
Model
N
Format ↑
Halluc. ↓
Missing ↓
Type acc. ↑
(%)
(%)
(%)
(%)
Gemma-4-E4B-it (vanilla)
33,977
100.0
6.7
75.4
60.8
Gemma-4-E4B-it (fine-tuned)
33,977
100.0
18.5
7.9
62.4
SELDNet + Gemma-4-E4B-it
33,977
100.0
50.1
20.0
56.3
SEA-LM (ours)
33,977
100.0
5.8
4.9
63.8
Appendix
Table 10: Complete Task 3 (Holistic Captioning) results on the evaluation set. Spatial metrics evaluate external sources; temporal and transcription metrics evaluate all applicable matched events. N = total evaluated examples; Format = format adherence; Halluc. = external-source hallucination rate; Missing = missing external source rate; Type acc. = source-type accuracy; Azim. = azimuth MAE; Elev. = elevation MAE; Dist. = distance MAE; Dist. err. = percentage distance error; Loc. bin = location-bin accuracy; Dist. bin = distance-bin accuracy; Start = start-time MAE; End = end-time MAE; tIoU = temporal intersection over union; WER = word error rate; Wearer acc. = example-level wearer-detection accuracy; Precision = precision; Recall = recall; F1 = F1 score.
Model
N
Format ↑
Halluc. ↓
Missing ↓
(%)
(%)
(%)
Gemma-4-E4B-it (vanilla)
25,876
100.0
15.8
45.0
Gemma-4-E4B-it (fine-tuned)
25,876
100.0
15.6
3.4
SELDNet + Gemma-4-E4B-it
25,876
100.0
10.9
16.1
SEA-LM (ours)
25,876
100.0
4.5
1.8
Appendix
Table 11: Complete Task 4 (Role-Aware Transcription) results on the evaluation set. Spatial metrics evaluate matched bystander speech; temporal and transcription metrics evaluate all applicable matched speech. N = total evaluated examples; Format = format adherence; Halluc. = bystander-speech hallucination rate; Missing = missing external source rate; Azim. = azimuth MAE; Elev. = elevation MAE; Dist. = distance MAE; Dist. err. = percentage distance error; Loc. bin = location-bin accuracy; Dist. bin = distance-bin accuracy; Start = start-time MAE; End = end-time MAE; tIoU = temporal intersection over union; WER = word error rate; Wearer acc. = example-level wearer-detection accuracy; Precision = precision; Recall = recall; F1 = F1 score.
Model
N
Format ↑
Start ↓
End ↓
tIoU ↑
WER ↓
Wearer acc. ↑
Precision ↑
Recall ↑
F1 ↑
(%)
(s)
(s)
(%)
(%)
(%)
(%)
(%)
Gemma-4-E4B-it (vanilla)
17,372
91.8
3.20
4.44
0.37
38.5
77.8
100.0
77.8
87.5
Gemma-4-E4B-it (fine-tuned)
17,372
100.0
0.05
0.34
0.95
8.9
100.0
100.0
100.0
100.0
SELDNet + Gemma-4-E4B-it
17,372
100.0
2.27
2.34
0.53
6.5
100.0
100.0
100.0
100.0
SEA-LM (ours)
17,372
100.0
0.02
0.04
0.98
5.8
100.0
100.0
100.0
100.0
Appendix
Table 12: Complete Task 5 (Targeted Wearer Speech Transcription) results on the evaluation set. N = total evaluated examples; Format = format adherence; Start = start-time MAE; End = end-time MAE; tIoU = temporal intersection over union; WER = word error rate; Wearer acc. = example-level wearer-detection accuracy; Precision = precision; Recall = recall; F1 = F1 score.
Model
N
Format ↑
Halluc. ↓
Missing ↓
(%)
(%)
(%)
Gemma-4-E4B-it (vanilla)
14,921
100.0
25.8
64.2
Gemma-4-E4B-it (fine-tuned)
14,921
100.0
27.2
0.5
SELDNet + Gemma-4-E4B-it
14,921
100.0
13.4
6.3
SEA-LM (ours)
14,921
100.0
8.0
1.2
Appendix
Table 13: Complete Task 6 (Targeted External Speech Transcription) results on the evaluation set. N = total evaluated examples; Format = format adherence; Halluc. = external-speech hallucination rate; Missing = missing external source rate; Azim. = azimuth MAE; Elev. = elevation MAE; Dist. = distance MAE; Dist. err. = percentage distance error; Loc. bin = location-bin accuracy; Dist. bin = distance-bin accuracy; Start = start-time MAE; End = end-time MAE; tIoU = temporal intersection over union; WER = word error rate; Detection acc. = source-level query-detection accuracy; Precision = precision; Recall = recall; F1 = F1 score.
Recent multimodal large language models mainly process audio as monaural signals, thereby discarding the spatial cues contained in spatial audio for sound localization, spatial relation reasoning, and spatial scene understanding. We propose Spatial-Omni, a lightweight method that implements SO-Encoder to inject First-Order Ambisonics (FOA) spatial audio into existing Omni LLMs as an independent modality, without modifying their original audio encoders. SO-Encoder provides spatial tokens with limited additional context cost and improves spatial audio understanding through efficient staged training. To support training and evaluation, we construct SO-Dataset, SO-QA, and SO-Bench from open-source data, real recordings, and simulations, containing 400K FOA spatial audio clips and 2.1M spatial question answering pairs. SO-Bench covers 16 spatial audio understanding subtasks, including basic detection and location estimation, spatial relation understanding, and complex spatial reasoning. Experiments show that Spatial-Omni outperforms existing open-source Large Audio-Language Models (LALMs) and Omni LLM models on spatial audio understanding tasks while retaining a reasonable level of general audio understanding. Code and data are available at https://github.com/dieKarotte/Spatial-Omni.
Spatial audio large language models (LLMs) enable embodied agents, wearable assistants, and immersive systems to recognize sound events, localize sources, and reason about their spatial relationships. However, existing spatial audio LLMs often rely on early fusion of acoustic and spatial features and source-agnostic token representations. These designs make it difficult to preserve the correspondence between individual sound events and their spatial attributes, particularly in multi-source scenes. To address this limitation, we propose SAIL, a Spatial Audio Intelligence framework with LLMs that preserves acoustic-spatial structure and source-level correspondence from audio encoding to LLM alignment. SAIL introduces a Disentangled Spatial Audio Transformer that represents Mel-spectrogram and interaural phase difference features as separate acoustic and spatial streams. Source-discriminative task queries further learn event, direction, and distance information for each source. A Dual-Stream Q-Former then aligns the two streams with the LLM using acoustic and spatial queries organized by source slots. Compared with the early-fusion baseline, SAIL achieves consistent improvements in dual-source sound event detection, direction and distance estimation, and spatial reasoning. These results demonstrate the importance of structured, source-discriminative audio representations for multi-source spatial understanding and reasoning.
Zhengding Luo, Jinyang Wu, Haozhe Ma +3
School of Electrical and Electronic Engineering, Nanyang Technological University, Singapore · Singapore Management University · Tencent Hy Frontier Lab, Singapore +2
Sound events are entities with semantic identities, locations, and trajectories, but current audio-language models usually reason about clips as global event content. Conversely, sound event localization models track source directions over time but offer limited semantic coverage for language reasoning. To address this gap, we introduce ST-AudioQA, a spatio-temporal audio QA dataset and benchmark built from first-order ambisonic (FOA) renderings of static and moving sound sources. Each scene provides source identity, activity, direction, distance, and motion metadata, enabling dense trajectory supervision and questions about what is sounding, where it is, how it moves, and how sources relate. We further propose ST-Audio Encoder, a time-resolved FOA audio encoder that learns event semantics together with source trajectories, and ST-AudioLM, which connects the audio tokens from the encoder to an LLM for spatio-temporal audio QA. Experiments show that this representation improves the semantic-localization tradeoff and yields stronger reasoning performance than static spatial and localization-oriented baselines.