RMS-AQA: A Two-Stage Spatial Audio Question Answering Benchmark for Real-World Domestic Environments
Authors: Peihao Chen, Qing Wang, Lichun Fan, Yufeng Hao, Zhifeng Kong, Mengyao Zhu, Hengyi Hong, Hang Chen, +7 more
Organizations: University of Science and Technology of China, China · Xiaomi Corporation, China · DataoceanAI, China · NVIDIA, USA · Soochow University, China
Embodied assistants in domestic environments must infer what happened, where and when it occurred, and how to respond. To address this, we introduce RMS-AQA, a spatial audio question answering (SAQA) benchmark for real-world domestic environments. The benchmark features a two-stage question-answering (QA) format to comprehensively assess the ability of audio-language models (ALMs) to first ground audible sound events and subsequently perform complex spatio-temporal reasoning based on that grounding. To maximize acoustic realism, our dataset combines authentic real-world first-order Ambisonics (FOA) recordings with high-fidelity synthetic data generated using measured room impulse responses (RIRs). Furthermore, we provide a lightweight spatial plug-in that injects FOA-format data into frozen audio-language backbones. Experimental results reveal that the primary challenges stem from concurrent sources, far distance, and sim-to-real domain gap between RIR-synthesized and authentic recordings.
Figures & tables
Figure 1: Overview of the RMS-AQA benchmark. (a) The data collection pipeline. (b) The six reasoning dimensions of the benchmark.
Dataset
# Clips
# QA pairs
# Rooms
Train
170K
170K
1,348
Validation
5K
6K
61
Test
2K
3K
5
Total
177K
179K
1,414
Table 1: Overview of the RMS-AQA dataset.
Figure 2: Baseline overview. The frozen semantic branch and the trainable spatial branch are fused for two-stage answer generation.
Model
Overall
Raw Stage-2 AS2
Grounded Stage-2 AGS2
AS1
AS2
AGS2
SC
SL
TD
SR
TR
AP
SC
SL
TD
SR
TR
AP
AF3 [ 8 ]
41.02
34.00
15.78
29.50
31.10
28.10
39.00
40.10
36.20
9.80
14.60
14.60
17.30
20.80
17.60
+ SpatialAug
63.98
52.60
39.10
59.10
36.40
48.90
45.40
57.30
68.50
46.40
25.20
36.90
29.80
44.30
52.00
AF-Next [ 7 ]
41.18
35.57
15.95
35.40
31.90
28.30
37.70
40.30
39.80
13.10
14.20
13.20
16.10
19.90
19.20
+ SpatialAug
65.93
60.23
45.17
69.30
40.80
62.70
47.90
62.20
78.50
54.50
29.00
46.40
33.40
49.30
58.40
MiDashengLM [ 2 ]
23.63
34.97
9.40
46.20
28.40
24.00
36.40
39.60
35.20
13.90
6.80
6.70
9.60
10.80
8.60
Table 2: Overall and category-wise accuracy (%) on the RMS-AQA validation set. AS1 , AS2 , and AGS2 denote Stage-1, raw Stage-2, and grounded Stage-2 accuracy, respectively. Rows marked “+ SpatialAug” report the adapted versions of the corresponding backbones.
Model
w/o Context
Predicted
Reference
AF3 + SpatialAug
44.12
52.60
60.05
Qwen3-Omni + SpatialAug
58.00
63.45
70.40
Table 3: Ablation on Stage-1 context for raw Stage-2 accuracy (%). “w/o Context” evaluates context-free reasoning. “Predicted” and “Reference” use generated and oracle Stage-1 answers, respectively.
Figure 3: Impact of scene complexity on grounded Stage-2 accuracy for the adapted Qwen3-Omni across simulated and recorded data.