cs.LGSep 30, 2026

Synchronous Multi-view Neural Diffusion

Authors: Yongquan Shi, Weijun Huang, Yueyang Pi, Wendi Zhao, Yiqing Shi, Shiping Wang

Abstract

Multi-view learning seeks to learn more comprehensive representations by exploiting the complementarity and consistency across diverse modalities or views. However, existing multi-view fusion strategies treat intra- and inter-view fusion as independent stages, without simultaneously considering the evolution within views and the dependency across views. Such an asynchronous fusion paradigm inevitably constrains cross-view interactions due to conflicting view-specific structural inductive biases. As a result, information flow is prone to distortion and compression along intermediate pathways, confining the model to learn within a restricted solution space. To address this, we propose Synchronous Multi-view Neural Diffusion (SynMDiff), which conceptualizes the multi-view feature space as a unified dynamical system driven by a diffusion process. By modeling the diffusion flow across arbitrary dyadic feature interactions in a joint space, SynMDiff enables the concurrent and adaptive intra- and inter-view information fusion. While a direct implementation of this synchronized mechanism incurs prohibitive computational costs, we further introduce an energy-based topological sampling strategy and an Ego-Net style centralized training architecture, ensuring both efficiency and scalability during learning and inference. Due to its conceptual elegance and computational efficacy, evaluations on real-world datasets demonstrate that SynMDiff outperforms the baselines by a large margin.

Figures & tables

Appendix figures & tables7 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. StreetDiff: Multi-view Street Scenes Generation via Cross-view Consistent Multi-view Stable Diffusion with Structure Prompts

    Sep 9, 2026Qi Zhang, Yanyifan Wang, Weiyuan Zhang +1Cross-ViewVideo Diffusion Models

  2. Divergence-Based Similarity Function for Multi-View Contrastive Learning

    Jul 9, 2025Jaehyoung Jeon, Cheolsu Lim, Myungjoo KangContrastive LearningMulti-View

  3. MV-Forcing: Long Multi-View Video Generation via 4D-Grounded Spatio-Temporal Self-Forcing

    Jul 6, 2026Gal Fiebelman, Hadar Averbuch-Elor, Sagie BenaimVideo Diffusion ModelsCamera-Conditioned Video Generation