eess.ASSep 30, 2026

MAV-C: A Framework for the Joint Objective Estimation of Audio-Visual Complexity in Immersive Virtual Environments

Authors: Luca Resti, Amelia Gully, Michael McLoughlin, Gavin Kearney, Alena Denisova

Organizations: Department of Computer Science, University of York, UK · AudioLab, School of Physics, Engineering and Technology, University of York, UK · Department of Language and Linguistic Science, University of York, UK

Abstract

We present the Motion-Aware Audio-Visual Complexity Metric (MAV-C), a reference-free framework for the joint objective estimation of audio-visual complexity. The metric combines entropy-based audio features (temporal, spectral, and spatial) with visual features (Sobel gradient magnitude, chromatic uniqueness, and optical flow) via a parametric fusion stage, producing a continuous joint complexity score CAV (t) [0,1]. We validate MAV-C on two datasets: a controlled synthetic corpus (SYN) of stimuli with known signal characteristics and a naturalistic gameplay corpus (GAM) of 60 clips drawn from the SAFEPLAY-X dataset. On SYN, the metric exhibits strong validity: the audio score CA and visual score CV are each insensitive to changes in the opposite modality (CoV < 0.003), the joint score CAV spans [0.00,0.90] across all parameter combinations, and single-axis feature sweeps produce monotone trajectories (Spearman up to 0.995). On GAM, CV differs significantly across content categories (Kruskal-Wallis p = 0.021) while CA does not, and the two sub-scores are uncorrelated (r = 0.03), confirming they operate on independent signal dimensions. OFAT sensitivity analysis identifies a two-tier parameter hierarchy, with modality balance (wa) and visual regularization (v) as most significant tunable parameters. Full subjective calibration is planned as future work.

Figures & tables

Explore similar work

CardsList
  1. Multimodal Confidence Modeling in Audio-Visual Quality Assessment

    May 2, 2026Mayesha Maliha R. Mithila, Mylene C. Q. FariasCross-Modal Attention

  2. AViS-Mamba: Adaptive Visual Steering of Audio State-Space Dynamics for Violence Detection

    Apr 2, 2026Damith Chamalke Senadeera, Dimitrios Kollias, Gregory SlabaughAudio-Visual ReasoningFine-Grained Video Understanding

  3. MultiRef-Compass: Towards Comprehensive Evaluation of Multi-Reference-to-Audio-Video Generation

    Jul 15, 2026Xiaohan Zhang, Yuqing Wen, Junlin Chen +9Audio-Video GenerationAudio-Visual Consistency