Multichannel Audio Quality Assessment: Extending Pretrained Perceptual Models to Spatial Audio
Organizations: Dolby Laboratories
Abstract
Accurate perceptual quality assessment is essential for evaluating and optimizing spatial audio, where perceived quality depends on both signal fidelity and inter-channel spatial relationships. However, subjective evaluation is costly, while existing perceptual models are often trained for limited channel configurations and cannot be directly applied to higher-channel-count audio. This raises the question: how can pretrained perceptual knowledge be effectively reused for multichannel spatial audio? Using 5.1-channel audio, we study four levels of multichannel integration: signal, prediction, latent, and feature and propose two learned approaches: latent-level aggregation of spatial-group representations and the feature-level Feature-Band Group Attention (FGAtt), which adaptively fuses spatial groups at the feature level before perceptual processing. Across five 5.1-channel test sets, FGAtt achieves the strongest over- all performance, demonstrating the effectiveness of feature-level adaptation for reusing pretrained perceptual knowledge
Figures & tables
| T1 | T2 | T3 | T4 | T5 | Overall | |||||||
| Method | PCC | SCC | PCC | SCC | PCC | SCC | PCC | SCC | PCC | SCC | PCC | SCC |
| Zimtohrli [ 1 ] | 0.7228 | 0.5907 | 0.9181 | 0.9138 | 0.8210 | 0.6948 | 0.8365 | 0.7245 | 0.8255 | 0.8129 | 0.7763 | 0.6872 |
| ITU Downmix [ 9 ] + PEAQ [ 15 ] | 0.5846 | 0.7686 | 0.7948 | 0.8753 | 0.5199 | 0.8200 | 0.3814 | 0.5376 | 0.3023 | 0.2182 | 0.4667 | 0.5944 |
| Binauralizer [ 2 ] + PEAQ [ 15 ] | 0.5684 | 0.7483 | 0.5970 | 0.6086 | 0.5216 | 0.8092 | 0.2969 | 0.4244 | 0.2187 | 0.0953 | 0.4051 | 0.5112 |
| ITU Downmix [ 9 ] + ViSQOL [ 7 ] | 0.8419 | 0.7428 | 0.8739 | 0.8643 | 0.8934 | 0.7480 | 0.8682 | 0.7720 | 0.8646 | 0.7125 | 0.8573 | 0.7455 |
| Binauralizer [ 2 ] + ViSQOL [ 7 ] | 0.8477 | 0.8791 | 0.8895 | 0.9204 | 0.8578 | 0.8798 | 0.8507 | 0.7810 | 0.8378 | 0.7549 | 0.8395 | 0.8066 |