cs.SDOct 1, 2026

RMS-AQA: A Two-Stage Spatial Audio Question Answering Benchmark for Real-World Domestic Environments

Authors: Peihao Chen, Qing Wang, Lichun Fan, Yufeng Hao, Zhifeng Kong, Mengyao Zhu, Hengyi Hong, Hang Chen, +7 more

Organizations: University of Science and Technology of China, China · Xiaomi Corporation, China · DataoceanAI, China · NVIDIA, USA · Soochow University, China

Abstract

Embodied assistants in domestic environments must infer what happened, where and when it occurred, and how to respond. To address this, we introduce RMS-AQA, a spatial audio question answering (SAQA) benchmark for real-world domestic environments. The benchmark features a two-stage question-answering (QA) format to comprehensively assess the ability of audio-language models (ALMs) to first ground audible sound events and subsequently perform complex spatio-temporal reasoning based on that grounding. To maximize acoustic realism, our dataset combines authentic real-world first-order Ambisonics (FOA) recordings with high-fidelity synthetic data generated using measured room impulse responses (RIRs). Furthermore, we provide a lightweight spatial plug-in that injects FOA-format data into frozen audio-language backbones. Experimental results reveal that the primary challenges stem from concurrent sources, far distance, and sim-to-real domain gap between RIR-synthesized and authentic recordings.

Figures & tables

Explore similar work

CardsList
  1. Spatio-Temporal Audio Language Modeling for Dynamic Sound Sources

    Jun 12, 2026Oh Hyun-Bin, Kazuki Shimada, Yuhta Takida +6Sound Source LocalizationLarge Audio Language Models

  2. Listen, See and Track: Spatio-Temporal Audio-Visual Sound Event Reasoning for Omni-Modal Language Models

    Aug 10, 2026Zhi Zeng, Cheng Zhang, Zesheng Yang +9Audio-Visual ReasoningSpatial Audio

  3. Spatial-Omni: Spatial Audio Understanding Integration in Multimodal LLMs via FOA Encoding

    Jun 9, 2026Zhiyuan Zhu, Yixuan Chen, Yiwen Shao +13Spatial AudioOmni-Modal