cs.CVSep 28, 2026

SyncRA: Learning Temporal Correspondence in Omni-Modal Models

Authors: Zelong Xu, Yan Li, Wenhe Hu, Xiyang Hu

Organizations: University of Wisconsin–Madison · University of Alberta · Arizona State University

Abstract

Recent omni-modal models demonstrate strong perception of audio and visual inputs, yet often struggle to connect what they hear with what they see at the same moment. This weakness in temporal correspondence can cause models to associate spoken cues with the wrong visual scenes, producing plausible answers grounded in incorrect audio-visual pairings. We diagnose this problem through controlled temporal swaps, revealing that model answers do not reliably follow changes in these pairings. To address it, we propose Synchrony-Guided Representation Alignment (SyncRA), a lightweight method for strengthening temporal correspondence between audio and vision. Specifically, SyncRA contrasts intermediate audio-visual representations within each video, aligning matching moments while separating mismatched ones to capture local temporal correspondence within a shared global context. The objective derives supervision directly from existing input timing, requiring no additional annotations and leaving inference unchanged. We evaluate SyncRA across four open omni-modal models spanning different sizes and architectures on five public video benchmarks. SyncRA consistently outperforms answer-only fine-tuning across all model-benchmark combinations, while substantially improving the ability to track changing audio-visual pairings in controlled evaluations. These results demonstrate that lightweight, targeted supervision can effectively strengthen temporal correspondence and translate into broad improvements in audio-visual question answering.

Figures & tables

Appendix figures & tables36 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. AVTrace: Diagnosing Audio-Visual Temporal Reasoning in Omni Models

    Sep 17, 2026Longyin Zhang, Parth Sakhare Mahendra, Chengwei Wei +4Audio-Visual ReasoningOmni-Modal

  2. Listen, See and Track: Spatio-Temporal Audio-Visual Sound Event Reasoning for Omni-Modal Language Models

    Aug 10, 2026Zhi Zeng, Cheng Zhang, Zesheng Yang +9Audio-Visual ReasoningSpatial Audio

  3. LatentOmni: Rethinking Omni-Modal Understanding via Unified Audio-Visual Latent Reasoning

    May 21, 2026Yifan Dai, Zhenhua Wu, Bohan Zeng +18Audio-Visual ReasoningOmni-Modal