cs.SDOct 7, 2026

SAVU-BENCH: A Real-World Benchmark for Spatial Audio-Visual Understanding

Authors: Yu Chen, Ruihang Liu, Yangguang Xu, Xinyue Jiang, Mohammed Bennamoun, Farid Boussaid, Xinyuan Qian, Qiuhong Ke

Abstract

Spatial audio-visual understanding requires models to recognize not only what is present, but also where events occur and how they relate across modalities. Existing benchmarks often rely on simulated scenes, evaluate isolated spatial skills, and provide limited diagnostic insight into failure modes. We introduce SAVU-Bench, a real-world benchmark that systematically evaluates spatial audio-visual understanding across three capability levels and seven evaluation tasks. We further introduce SAVU-Diag, a scene-linked diagnostic set that decomposes reasoning questions into their prerequisite grounding and alignment sub-tasks. Evaluation of 12 representative models on SAVU-Bench reveals that while visual spatial grounding is relatively mature, spatial perception involving audio remains a primary bottleneck. SAVU-Diag further demonstrates that most reasoning errors co-occur with failures on these prerequisite tasks, though reasoning gaps persist even when prerequisites are correctly resolved. Motivated by these findings, we introduce SAVU-EA, a training-free evidence-augmented baseline that makes spatial cues more explicit. While SAVU-EA substantially improves spatial grounding and joint matching, high-level spatial reasoning remains challenging. Our findings highlight the urgent need for both robust spatial audio perception and deeper integration of cross-modal spatial relations.

Explore similar work

CardsList
  1. FloorSAV: Elucidating Spatial Audio-Visual Context with 2D Floormap for AV-LLMs

    Oct 8, 2026Kyeong-Rae Kim, Sungnyun Kim, Tae-Hyun Oh3D Spatial ReasoningAudio-Visual Understanding

  2. OmniEcho: Audio-Visual Spatial Understanding for Omni-Modal Embodied Agents

    Sep 20, 2026Ruixun Liu, Yuxuan Wang, Jiacheng Xie +10Audio-Visual UnderstandingSpatial Reasoning Benchmarks

  3. Unlocking Spatial Grounding in Large Audio-Visual Retrieval models

    Jun 22, 2026Hugo Malard, Michel Olvera, Sanjeel Parekh +3Multimodal GroundingAudio-Visual Understanding