Large audio-language models (LALMs) have achieved strong performance in general audio understanding, yet most are designed for monaural input and discard the inter-channel cues essential for spatial perception. In contrast, existing spatial audio-language models are purpose-built for spatial tasks and fail to capitalize on the general understanding capabilities of monaural LALMs. We present MiDashengLM-Spatial, the first open-source end-to-end unified audio-language model, to our knowledge, which supports both general audio understanding and spatial awareness within a single architecture. It extends MiDashengLM with a spatial audio encoder, Spatial-Dasheng, integrated through a hierarchical semantic-to-spatial conditioning module that injects intermediate semantic representations into the spatial branch at multiple depths while preserving the original semantic pathway. To provide spatial audio-language supervision at scale, we develop a data synthesis pipeline that renders diverse spatial acoustic scenes with scene-level spatial descriptions and question-answer pairs. Experiments show that Spatial-Dasheng achieves strong performance on sound event localization and detection in real-world scenes, and that MiDashengLM-Spatial substantially outperforms existing LALMs on spatial understanding and reasoning benchmarks. Meanwhile, it remains competitive with state-of-the-art 8B-scale LALMs on diverse monaural benchmarks, demonstrating that spatial awareness can be acquired without compromising general audio understanding. The source code and model checkpoint are available at https://github.com/xiaomi-research/midashenglm-spatial and https://huggingface.co/mispeech/midashenglm-spatial.
Figures & tables
Figure 1: An example of a typical spatial acoustic scene. Left: a top-down view of a room, illustrating the categories, positions, and movement trajectories of sound events. Right: the timeline of the corresponding sound events.
Figure 2: The overall architecture of MiDashengLM-Spatial. The left part shows the details of the dual-branch audio encoder.
Appendix figures & tables4 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 3: Mapping from spatial attributes to textual descriptions. (a) azimuth sectors (top view) and (b) elevation bands (side view), both defined in listener-centered coordinates. Sources with ∣ϕ∣≤5∘ are treated as horizontal and described without the up or down term.
Category
Question-Answer Template
Source Localization
Source-to-Direction Localization
Q : From which direction is {sound} coming? / From which direction do you hear the person who says {speech} ? A : {dir} Q : Is {sound} above or below the listener? A : above / below / Level with the listener
Direction-to-Source Identification
Q : Which sound comes from the {dir} ? A : {sound} Q : What does the person on your {dir} say? A : {speech}
Motion Trajectory Perception
Trajectory Tracking
Q : Does {sound} / the person who says {speech} move or stay in place? A : It moves / It stays in place / It cannot be determined Q : How does {sound} / the person who says {speech} move? A : Clockwise / Counter-clockwise / Stays in place Q : Where does {sound} / the person who says {speech} start out / end up? A : {dir} (the start/end position of the movement path)
Distance Change Tracking
Q : Does {sound} / the person who says {speech} move closer to or farther from the listener (you)? A : Getting closer / Moving away / Distance unchanged
Appendix
Table 4: Categories and question-answer templates of the deterministically constructed spatial QA pairs. Sound-scene and speech-scene QA share the same templates. {sound} : a short description of a sound source (e.g., a dog barking ); {speech} : a short summary of what a speaker says; {dir} : a quantized direction phrase (e.g., front-left ).
Figure 4: Examples of reference and predicted outputs for spatial audio captioning and spatial speech transcription, together with the structured attributes reverse-parsed into JSON format by an LLM.
Recent multimodal large language models mainly process audio as monaural signals, thereby discarding the spatial cues contained in spatial audio for sound localization, spatial relation reasoning, and spatial scene understanding. We propose Spatial-Omni, a lightweight method that implements SO-Encoder to inject First-Order Ambisonics (FOA) spatial audio into existing Omni LLMs as an independent modality, without modifying their original audio encoders. SO-Encoder provides spatial tokens with limited additional context cost and improves spatial audio understanding through efficient staged training. To support training and evaluation, we construct SO-Dataset, SO-QA, and SO-Bench from open-source data, real recordings, and simulations, containing 400K FOA spatial audio clips and 2.1M spatial question answering pairs. SO-Bench covers 16 spatial audio understanding subtasks, including basic detection and location estimation, spatial relation understanding, and complex spatial reasoning. Experiments show that Spatial-Omni outperforms existing open-source Large Audio-Language Models (LALMs) and Omni LLM models on spatial audio understanding tasks while retaining a reasonable level of general audio understanding. Code and data are available at https://github.com/dieKarotte/Spatial-Omni.
Embodied, ego-centric intelligence fundamentally requires the ability to comprehend spatial audio within complex environments. While large audio-language models excel at mono-channel reasoning, they lack spatial awareness, discarding critical spatial cues that enable sound localization and that can improve the disentanglement of overlapping sound sources. To address this, we present SEA-LM, a Spatial Audio Understanding model. First, we introduce FOACODER, a layout-flexible spatial audio encoder trained on source localization and ego-centric voice activity detection objectives to encode First Order Ambisonics derived from variable-count, variable-position smart-glasses arrays via beamforming. We train a Multimodal Large Language Model (MLLM) to understand these spatial audio embeddings through a two-stage curriculum spanning six tasks, including sound localization and spatially selective transcription in settings with multiple speakers and overlapping sounds. To prevent the transcription outputs from dominating the next token prediction loss and overwhelming the direction predictions, we introduce a Spatio-temporal Weighted Cross-Entropy Loss. On our evaluation set, SEA-LM achieves lower azimuth and elevation MAE, higher temporal IoU, lower external-source hallucination and missing-source rates, and lower WER on most transcription tasks than compared baselines, while remaining robust across 1,211 smart-glasses array configurations with 4 to 9 microphones.
Sonal Kumar, Sinan Hersek, Artem Dementyev +6
University of Maryland, College Park · Google · Google DeepMind
Spatial audio large language models (LLMs) enable embodied agents, wearable assistants, and immersive systems to recognize sound events, localize sources, and reason about their spatial relationships. However, existing spatial audio LLMs often rely on early fusion of acoustic and spatial features and source-agnostic token representations. These designs make it difficult to preserve the correspondence between individual sound events and their spatial attributes, particularly in multi-source scenes. To address this limitation, we propose SAIL, a Spatial Audio Intelligence framework with LLMs that preserves acoustic-spatial structure and source-level correspondence from audio encoding to LLM alignment. SAIL introduces a Disentangled Spatial Audio Transformer that represents Mel-spectrogram and interaural phase difference features as separate acoustic and spatial streams. Source-discriminative task queries further learn event, direction, and distance information for each source. A Dual-Stream Q-Former then aligns the two streams with the LLM using acoustic and spatial queries organized by source slots. Compared with the early-fusion baseline, SAIL achieves consistent improvements in dual-source sound event detection, direction and distance estimation, and spatial reasoning. These results demonstrate the importance of structured, source-discriminative audio representations for multi-source spatial understanding and reasoning.
Zhengding Luo, Jinyang Wu, Haozhe Ma +3
School of Electrical and Electronic Engineering, Nanyang Technological University, Singapore · Singapore Management University · Tencent Hy Frontier Lab, Singapore +2