Accurate perceptual quality assessment is essential for evaluating and optimizing spatial audio, where perceived quality depends on both signal fidelity and inter-channel spatial relationships. However, subjective evaluation is costly, while existing perceptual models are often trained for limited channel configurations and cannot be directly applied to higher-channel-count audio. This raises the question: how can pretrained perceptual knowledge be effectively reused for multichannel spatial audio? Using 5.1-channel audio, we study four levels of multichannel integration: signal, prediction, latent, and feature and propose two learned approaches: latent-level aggregation of spatial-group representations and the feature-level Feature-Band Group Attention (FGAtt), which adaptively fuses spatial groups at the feature level before perceptual processing. Across five 5.1-channel test sets, FGAtt achieves the strongest over- all performance, demonstrating the effectiveness of feature-level adaptation for reusing pretrained perceptual knowledge
Figures & tables
T1
T2
T3
T4
T5
Overall
Method
PCC
SCC
PCC
SCC
PCC
SCC
PCC
SCC
PCC
SCC
PCC
SCC
Zimtohrli [ 1 ]
0.7228
0.5907
0.9181
0.9138
0.8210
0.6948
0.8365
0.7245
0.8255
0.8129
0.7763
0.6872
ITU Downmix [ 9 ] + PEAQ [ 15 ]
0.5846
0.7686
0.7948
0.8753
0.5199
0.8200
0.3814
0.5376
0.3023
0.2182
0.4667
0.5944
Binauralizer [ 2 ] + PEAQ [ 15 ]
0.5684
0.7483
0.5970
0.6086
0.5216
0.8092
0.2969
0.4244
0.2187
0.0953
0.4051
0.5112
ITU Downmix [ 9 ] + ViSQOL [ 7 ]
0.8419
0.7428
0.8739
0.8643
0.8934
0.7480
0.8682
0.7720
0.8646
0.7125
0.8573
0.7455
Binauralizer [ 2 ] + ViSQOL [ 7 ]
0.8477
0.8791
0.8895
0.9204
0.8578
0.8798
0.8507
0.7810
0.8378
0.7549
0.8395
0.8066
Table 1: Perceptual quality prediction performance on five 5.1-channel test sets. LFE is discarded. GMLv2 is frozen. Zimtohrli evaluates each channel independently, with predictions averaged across channels.
Figure 1: Overall correlation between predicted and ground-truth MUSHRA scores across GMLv2 spatial aggregation strategies. The dashed line denotes the identity function.
Pretrained spatial audio encoders are increasingly used as general-purpose representations for perceptual tasks, yet their spatial encoding capabilities remain poorly understood. We introduce the Spatial Audio Representation Learning (SARL) benchmark, a controlled framework for evaluating spatial information in pretrained audio models. SARL probes source-level factors (azimuth, elevation, distance, class) and room-level factors (RT60, volume, shape). Experiments across diverse encoders reveal three patterns: input configuration and training paradigm shape spatial encoding; source factors are consistently easier to decode than room factors; and sensitivity analysis under controlled perturbations shows heterogeneous responses to source and room variation. These results reveal systematic biases in current pretrained audio representations. SARL is released as an open-source benchmark for reproducible evaluation of spatial audio representations.
Chuyang Chen, Sivan Ding, Adrian S. Roman +1
Music and Audio Research Laboratory, New York University, USA
Multimodal Large Language Models (LLMs) have remarkable semantic audio understanding, yet they remain "spatially agnostic" due to their reliance on mono-channel audio representations. Currently, spatial audio perception methods mainly focus on complex room simulations and custom-trained, geometry-aware stereo encoders, which limits their accessibility and generalizability. In this paper, we introduce the Dual-BEATs architecture, in which the left and right audio channels are routed independently through two identical semantic encoders as an alternative to specialized spatial modules. To circumvent the architectural bottleneck where internal normalization otherwise erases the inter-channel variance of stereo audio, we inject a static, uncorrelated dithering noise floor prior to encoding. This dithering intervention establishes a macro-variance floor that "smuggles" spatial geometry across the normalization layers. Evaluated on a ternary directional classification task (Left, Center, Right), we demonstrate that dithered models achieve exceptional spatial resolution--reaching up to 97.2% localization accuracy even on subtle 0.5 panning amplitudes--and demonstrates robust, zero-shot generalization to entirely unseen spatial configurations. Our results suggest that with the appropriate acoustic regularization, standard multimodal models are natively capable of generalized stereo audio understanding.
Shuo-Chun Lin, Hen-Hsen Huang
Institute of Information Science, Academia Sinica, Taiwan
Spatial audio large language models (LLMs) enable embodied agents, wearable assistants, and immersive systems to recognize sound events, localize sources, and reason about their spatial relationships. However, existing spatial audio LLMs often rely on early fusion of acoustic and spatial features and source-agnostic token representations. These designs make it difficult to preserve the correspondence between individual sound events and their spatial attributes, particularly in multi-source scenes. To address this limitation, we propose SAIL, a Spatial Audio Intelligence framework with LLMs that preserves acoustic-spatial structure and source-level correspondence from audio encoding to LLM alignment. SAIL introduces a Disentangled Spatial Audio Transformer that represents Mel-spectrogram and interaural phase difference features as separate acoustic and spatial streams. Source-discriminative task queries further learn event, direction, and distance information for each source. A Dual-Stream Q-Former then aligns the two streams with the LLM using acoustic and spatial queries organized by source slots. Compared with the early-fusion baseline, SAIL achieves consistent improvements in dual-source sound event detection, direction and distance estimation, and spatial reasoning. These results demonstrate the importance of structured, source-discriminative audio representations for multi-source spatial understanding and reasoning.
Zhengding Luo, Jinyang Wu, Haozhe Ma +3
School of Electrical and Electronic Engineering, Nanyang Technological University, Singapore · Singapore Management University · Tencent Hy Frontier Lab, Singapore +2