eess.ASSep 30, 2026

MAV-C: A Framework for the Joint Objective Estimation of Audio-Visual Complexity in Immersive Virtual Environments

Authors: Luca Resti, Amelia Gully, Michael McLoughlin, Gavin Kearney, Alena Denisova

Organizations: Department of Computer Science, University of York, UK · AudioLab, School of Physics, Engineering and Technology, University of York, UK · Department of Language and Linguistic Science, University of York, UK

Abstract

We present the Motion-Aware Audio-Visual Complexity Metric (MAV-C), a reference-free framework for the joint objective estimation of audio-visual complexity. The metric combines entropy-based audio features (temporal, spectral, and spatial) with visual features (Sobel gradient magnitude, chromatic uniqueness, and optical flow) via a parametric fusion stage, producing a continuous joint complexity score CAV (t) [0,1]. We validate MAV-C on two datasets: a controlled synthetic corpus (SYN) of stimuli with known signal characteristics and a naturalistic gameplay corpus (GAM) of 60 clips drawn from the SAFEPLAY-X dataset. On SYN, the metric exhibits strong validity: the audio score CA and visual score CV are each insensitive to changes in the opposite modality (CoV < 0.003), the joint score CAV spans [0.00,0.90] across all parameter combinations, and single-axis feature sweeps produce monotone trajectories (Spearman up to 0.995). On GAM, CV differs significantly across content categories (Kruskal-Wallis p = 0.021) while CA does not, and the two sub-scores are uncorrelated (r = 0.03), confirming they operate on independent signal dimensions. OFAT sensitivity analysis identifies a two-tier parameter hierarchy, with modality balance (wa) and visual regularization (v) as most significant tunable parameters. Full subjective calibration is planned as future work.

Figures & tables

Explore similar work

May 2, 2026cs.MM

Multimodal Confidence Modeling in Audio-Visual Quality Assessment

Audio-visual quality assessment (AVQA) is essential for streaming, teleconferencing, and immersive media. In realistic streaming scenarios, distortions are often asymmetric, where one modality may be severely degraded while the other remains clean. Still, most contemporary AVQA metrics treat audio and video as equally reliable, causing confidence-unaware fusion to emphasize unreliable signals. This paper proposes MCM-AVQA, a multimodal confidence-aware AVQA framework that explicitly estimates modality-specific confidence and injects it into a dedicated audio-visual mixer for cross-modal attention. The Audio-Visual Mixer utilizes frame-level, confidence-guided channel attention to gate fusion, modulating feature interaction between modalities so that high-confidence streams dominate while unreliable inputs are suppressed, preserving temporal degradation patterns. A multi-head visual confidence estimator turns frame-level artifact probabilities into temporally smoothed, clip-level visual confidence scores, while an audio confidence module derives confidence from speech-quality cues without requiring a clean reference. Experiments on multiple AVQA benchmarks show that MCM-AVQA, and specifically its confidence-guided Audio-Visual Mixer, improve correlation with human mean opinion scores and yield more interpretable behavior under real-world asymmetric audio-visual distortions.
Apr 2, 2026cs.CV

AViS-Mamba: Adaptive Visual Steering of Audio State-Space Dynamics for Violence Detection

Automatic violence detection from video is challenging because violent interactions may be distant, occluded, or only partially visible. Audio can provide complementary evidence for violent events that are difficult to recognize from visual information alone. However, audio itself may be absent, dubbed, or dominated by environmental noise, making the central challenge not whether to incorporate audio but how to adapt reliance on it according to the visual scene. We introduce \emph{AViS-Mamba}, an audiovisual Mamba-based architecture in which the visual stream directly governs the behavior of the audio stream. At each layer of the audio encoder, a compact visual representation produces a modulation vector that conditions the encoder's internal temporal operators together with a routing gate that regulates the strength of this visual intervention. Rather than fusing or reweighting features after they have been extracted, visual context directly shapes the temporal dynamics of the audio encoder. We further propose Adaptive AV-InfoNCE, a contrastive objective that learns to balance the audio-to-video and video-to-audio alignment directions rather than weighting them uniformly. On the audio-valid NTU-CCTV and DVD benchmarks, AViS-Mamba establishes state-of-the-art results, attaining 88.59% and 75.74% accuracy. We demonstrate that adaptive visual conditioning consistently outperforms fixed routing and improves performance under degraded and missing-audio conditions. Layer-wise analysis further reveals that the model adapts the audio stream selectively across network depth rather than applying a single global routing policy.
Jul 15, 2026cs.CV

MultiRef-Compass: Towards Comprehensive Evaluation of Multi-Reference-to-Audio-Video Generation

Multi-reference-to-audio-video (MR2AV) generation aims to generate coherent audio-video content conditioned on multiple references and textual instructions. Existing benchmarks mainly focus on text-driven generation, single-reference subject preservation, or isolated audio-video alignment, leaving the emerging MR2AV setting largely unexplored. Compared with these settings, MR2AV requires models to jointly reason over multiple references while generating synchronized visual and audio content. Models must not only preserve each reference faithfully but also correctly bind and compose multiple referenced entities into coherent audio-visual events. To address this gap, we introduce MultiRef-Compass, a unified benchmark for MR2AV generation. It comprises 350350 carefully curated samples constructed through a scalable and controllable asset-composition pipeline, covering multi-view subject preservation, multi-entity binding, and human-object-scene composition. To provide interpretable assessment, MultiRef-Compass defines an evaluation protocol with four dimensions: Basic Quality, Reference Consistency, Audio-Visual Consistency, and Instruction Following, using 14 sub-metrics. MultiRef-Compass integrates automatic metrics with a rejudging-enhanced MLLM-as-a-Judge framework, enabling scalable and auditable evaluation of both perceptual fidelity and reference-conditioned composition. Extensive experiments on eight representative MR2AV systems reveal substantial room for improvement across multiple evaluation dimensions, underscoring the need for a comprehensive benchmark and positioning MultiRef-Compass as a foundation for future MR2AV research.