RMS-AQA: A Two-Stage Spatial Audio Question Answering Benchmark for Real-World Domestic Environments
Authors: Peihao Chen, Qing Wang, Lichun Fan, Yufeng Hao, Zhifeng Kong, Mengyao Zhu, Hengyi Hong, Hang Chen, +7 more
Organizations: University of Science and Technology of China, China · Xiaomi Corporation, China · DataoceanAI, China · NVIDIA, USA · Soochow University, China
Embodied assistants in domestic environments must infer what happened, where and when it occurred, and how to respond. To address this, we introduce RMS-AQA, a spatial audio question answering (SAQA) benchmark for real-world domestic environments. The benchmark features a two-stage question-answering (QA) format to comprehensively assess the ability of audio-language models (ALMs) to first ground audible sound events and subsequently perform complex spatio-temporal reasoning based on that grounding. To maximize acoustic realism, our dataset combines authentic real-world first-order Ambisonics (FOA) recordings with high-fidelity synthetic data generated using measured room impulse responses (RIRs). Furthermore, we provide a lightweight spatial plug-in that injects FOA-format data into frozen audio-language backbones. Experimental results reveal that the primary challenges stem from concurrent sources, far distance, and sim-to-real domain gap between RIR-synthesized and authentic recordings.
Figures & tables
Figure 1: Overview of the RMS-AQA benchmark. (a) The data collection pipeline. (b) The six reasoning dimensions of the benchmark.
Dataset
# Clips
# QA pairs
# Rooms
Train
170K
170K
1,348
Validation
5K
6K
61
Test
2K
3K
5
Total
177K
179K
1,414
Table 1: Overview of the RMS-AQA dataset.
Figure 2: Baseline overview. The frozen semantic branch and the trainable spatial branch are fused for two-stage answer generation.
Model
Overall
Raw Stage-2 AS2
Grounded Stage-2 AGS2
AS1
AS2
AGS2
SC
SL
TD
SR
TR
AP
SC
SL
TD
SR
TR
AP
AF3 [ 8 ]
41.02
34.00
15.78
29.50
31.10
28.10
39.00
40.10
36.20
9.80
14.60
14.60
17.30
20.80
17.60
+ SpatialAug
63.98
52.60
39.10
59.10
36.40
48.90
45.40
57.30
68.50
46.40
25.20
36.90
29.80
44.30
52.00
AF-Next [ 7 ]
41.18
35.57
15.95
35.40
31.90
28.30
37.70
40.30
39.80
13.10
14.20
13.20
16.10
19.90
19.20
+ SpatialAug
65.93
60.23
45.17
69.30
40.80
62.70
47.90
62.20
78.50
54.50
29.00
46.40
33.40
49.30
58.40
MiDashengLM [ 2 ]
23.63
34.97
9.40
46.20
28.40
24.00
36.40
39.60
35.20
13.90
6.80
6.70
9.60
10.80
8.60
Table 2: Overall and category-wise accuracy (%) on the RMS-AQA validation set. AS1 , AS2 , and AGS2 denote Stage-1, raw Stage-2, and grounded Stage-2 accuracy, respectively. Rows marked “+ SpatialAug” report the adapted versions of the corresponding backbones.
Model
w/o Context
Predicted
Reference
AF3 + SpatialAug
44.12
52.60
60.05
Qwen3-Omni + SpatialAug
58.00
63.45
70.40
Table 3: Ablation on Stage-1 context for raw Stage-2 accuracy (%). “w/o Context” evaluates context-free reasoning. “Predicted” and “Reference” use generated and oracle Stage-1 answers, respectively.
Figure 3: Impact of scene complexity on grounded Stage-2 accuracy for the adapted Qwen3-Omni across simulated and recorded data.
Sound events are entities with semantic identities, locations, and trajectories, but current audio-language models usually reason about clips as global event content. Conversely, sound event localization models track source directions over time but offer limited semantic coverage for language reasoning. To address this gap, we introduce ST-AudioQA, a spatio-temporal audio QA dataset and benchmark built from first-order ambisonic (FOA) renderings of static and moving sound sources. Each scene provides source identity, activity, direction, distance, and motion metadata, enabling dense trajectory supervision and questions about what is sounding, where it is, how it moves, and how sources relate. We further propose ST-Audio Encoder, a time-resolved FOA audio encoder that learns event semantics together with source trajectories, and ST-AudioLM, which connects the audio tokens from the encoder to an LLM for spatio-temporal audio QA. Experiments show that this representation improves the semantic-localization tradeoff and yields stronger reasoning performance than static spatial and localization-oriented baselines.
Understanding dynamic sound sources requires jointly determining what produces a sound, where the source is located, and how it moves over time. Yet existing audio-language models often represent clips as global acoustic events, while vision-language models lack the spatial audio cues needed to localize and track individual sources. To evaluate this missing capability, we introduce ST-OmniQA, a spatio-temporal audio-visual question-answering benchmark built from panoramic videos paired with synchronized first-order Ambisonics (FOA) audio of moving sound sources. It contains 40K videos and 400K question-answer pairs organized into four capability levels covering sound-event recognition, direction of arrival, source distance, motion trajectories, and temporally grounded audio-visual reasoning. Building on this benchmark, we propose ST-Omni-R1, which integrates FOA-derived semantic and trajectory representations with panoramic visual context and is trained through progressive curriculum learning and reasoning-tree reinforcement learning. ST-Omni-R1 achieves 77.83% average semantic accuracy across the four levels, compared with 37.28% for the best evaluated baseline. Results on three public spatial-audio benchmarks further indicate that its learned spatial and motion representations transfer beyond ST-OmniQA.
Zhi Zeng, Cheng Zhang, Zesheng Yang +9
Xi’an Jiaotong University · The Hong Kong Polytechnic University · National University of Singapore +2
Recent multimodal large language models mainly process audio as monaural signals, thereby discarding the spatial cues contained in spatial audio for sound localization, spatial relation reasoning, and spatial scene understanding. We propose Spatial-Omni, a lightweight method that implements SO-Encoder to inject First-Order Ambisonics (FOA) spatial audio into existing Omni LLMs as an independent modality, without modifying their original audio encoders. SO-Encoder provides spatial tokens with limited additional context cost and improves spatial audio understanding through efficient staged training. To support training and evaluation, we construct SO-Dataset, SO-QA, and SO-Bench from open-source data, real recordings, and simulations, containing 400K FOA spatial audio clips and 2.1M spatial question answering pairs. SO-Bench covers 16 spatial audio understanding subtasks, including basic detection and location estimation, spatial relation understanding, and complex spatial reasoning. Experiments show that Spatial-Omni outperforms existing open-source Large Audio-Language Models (LALMs) and Omni LLM models on spatial audio understanding tasks while retaining a reasonable level of general audio understanding. Code and data are available at https://github.com/dieKarotte/Spatial-Omni.