Multi-view learning seeks to learn more comprehensive representations by exploiting the complementarity and consistency across diverse modalities or views. However, existing multi-view fusion strategies treat intra- and inter-view fusion as independent stages, without simultaneously considering the evolution within views and the dependency across views. Such an asynchronous fusion paradigm inevitably constrains cross-view interactions due to conflicting view-specific structural inductive biases. As a result, information flow is prone to distortion and compression along intermediate pathways, confining the model to learn within a restricted solution space. To address this, we propose Synchronous Multi-view Neural Diffusion (SynMDiff), which conceptualizes the multi-view feature space as a unified dynamical system driven by a diffusion process. By modeling the diffusion flow across arbitrary dyadic feature interactions in a joint space, SynMDiff enables the concurrent and adaptive intra- and inter-view information fusion. While a direct implementation of this synchronized mechanism incurs prohibitive computational costs, we further introduce an energy-based topological sampling strategy and an Ego-Net style centralized training architecture, ensuring both efficiency and scalability during learning and inference. Due to its conceptual elegance and computational efficacy, evaluations on real-world datasets demonstrate that SynMDiff outperforms the baselines by a large margin.
Figures & tables
Figure 1. Comparison of two multi-view fusion paradigms. (a) The two-stage asynchronous paradigm does not explicitly model cross-view interactions. Conflicts among view-specific structural inductive biases can therefore distort information flow along intermediate pathways. (b) The proposed synchronous paradigm jointly models intra-view and inter-view fusion, enabling direct interactions between arbitrary views. A two-panel schematic comparing asynchronous and synchronous multi-view fusion. Panel (a) shows view-specific information being processed separately and then fused in two stages. Indirect cross-view pathways contain conflicts caused by incompatible structural inductive biases, resulting in distorted information flow. Panel (b) shows the proposed synchronous paradigm, which jointly models intra-view and inter-view interactions and provides direct information-flow pathways among arbitrary views.
Method
BRCA
LGG
UCEC
GBMLGG
TCGA
Macro-F1
Micro-F1
Macro-F1
Micro-F1
Macro-F1
Micro-F1
Macro-F1
Micro-F1
Macro-F1
Micro-F1
SVM
43.6 ± 0.0
70.2 ± 0.0
60.0 ± 0.0
60.0 ± 0.0
28.0 ± 0.0
72.5 ± 0.0
39.4 ± 0.0
52.0 ± 0.0
61.4 ± 0.0
68.0 ± 0.0
RF
68.8 ± 0.0
80.2 ± 0.0
49.6 ± 0.0
49.8 ± 0.0
28.7 ± 0.0
72.5 ± 0.0
47.0 ± 0.0
54.6 ± 0.0
52.0 ± 0.0
67.0 ± 0.0
DeepMO
76.4 ± 4.9
81.4 ± 2.1
63.7 ± 3.8
64.2 ± 3.1
53.8 ± 1.9
80.2 ± 3.1
51.1 ± 2.2
53.5 ± 2.4
64.6 ± 6.9
73.0 ± 6.9
MOGONET
58.9 ± 2.6
71.6 ± 1.5
61.8 ± 2.4
62.3 ± 1.8
43.7 ± 0.6
75.4 ± 1.9
43.6 ± 2.7
49.2 ± 1.5
38.5 ± 0.4
42.1 ± 0.7
MoGCN
55.3 ± 0.7
73.6 ± 0.4
33.8 ± 0.0
51.1 ± 0.0
28.0 ± 0.0
71.6 ± 0.0
41.5 ± 8.4
53.0 ± 6.3
66.7 ± 0.4
73.9 ± 0.4
Table 1. Classification results (mean% ± std%) on cancer subtype datasets. Best and second-best results are highlighted in red and blue, respectively. “OOM” indicates out-of-memory errors.
Method
FreeBase
DBLP
IMDB
Yelp
AMiner
Macro-F1
Micro-F1
Macro-F1
Micro-F1
Macro-F1
Micro-F1
Macro-F1
Micro-F1
Macro-F1
Micro-F1
GCN
44.6 ± 1.4
37.4 ± 1.1
90.1 ± 0.8
91.6 ± 0.6
24.3 ± 0.2
55.4 ± 0.2
52.0 ± 0.2
67.4 ± 0.9
68.4 ± 0.6
81.6 ± 0.7
HAN
62.1 ± 2.4
48.8 ± 3.4
89.3 ± 0.4
90.4 ± 0.4
23.9 ± 0.5
55.9 ± 0.8
48.3 ± 0.3
48.9 ± 0.6
72.3 ± 0.6
84.8 ± 0.1
DMGI
54.8 ± 2.1
41.1 ± 1.9
65.7 ± 0.2
71.1 ± 1.0
35.3 ± 1.0
57.3 ± 0.8
51.6 ± 0.4
69.8 ± 0.2
30.3 ± 0.7
65.5 ± 0.5
IGNN
65.1 ± 0.1
61.7 ± 0.2
86.8 ± 0.1
87.5 ± 0.9
45.3 ± 0.3
54.8 ± 0.7
64.5 ± 0.4
71.2 ± 0.6
74.5 ± 0.6
85.2 ± 0.3
MRGCN
57.0 ± 0.3
53.9 ± 0.1
89.5 ± 0.3
90.5 ± 0.6
45.2 ± 0.6
47.7 ± 0.7
54.3 ± 0.4
73.7 ± 0.4
73.4 ± 0.4
82.9 ± 0.4
Table 2. Node classification results (mean% ± std%) on heterogeneous graph datasets. Best and second-best results are highlighted in red and blue, respectively.
Metrics
Methods
Co-GCN
PDMF
LGCN-FF
ECMGD
TUNED
KAMSSM
CoGFormer
SynMDiff
Macro-F1
Scene15
18.9 ± 8.7
39.8 ± 4.6
42.3 ± 5.7
69.3 ± 4.3
70.0 ± 3.0
67.8 ± 0.8
74.0 ± 0.7
81.1 ± 0.3
YouTube
43.4 ± 4.0
36.9 ± 3.3
42.3 ± 5.7
59.0 ± 0.4
57.3 ± 0.9
56.5 ± 1.1
63.6 ± 1.4
70.5 ± 0.5
MITIndoor
51.8 ± 0.8
48.9 ± 0.3
21.1 ± 7.8
36.5 ± 8.1
22.4 ± 3.4
25.8 ± 3.0
51.2 ± 2.3
55.2 ± 0.8
HW
94.9 ± 2.0
90.0 ± 2.3
91.5 ± 2.8
95.4 ± 0.0
88.9 ± 1.6
96.3 ± 0.2
96.6 ± 0.2
97.2 ± 0.2
IAPR
55.8 ± 3.8
60.1 ± 0.7
57.0 ± 1.4
65.5 ± 0.2
64.1 ± 4.4
65.9 ± 0.2
66.2 ± 0.2
70.2 ± 0.1
Animals
61.9 ± 4.3
70.6 ± 0.3
62.9 ± 6.2
74.8 ± 0.4
74.7 ± 0.6
70.9 ± 0.6
77.3 ± 0.2
79.1 ± 0.2
Table 3. Classification results (mean% ± std%) on multi-view datasets. Best and second-best results are highlighted in red and blue, respectively.
Dataset
NoisyMNIST
YTF
CIFAR-10
VGGFace
Samples
70,000
286,006
50,000
34,027
Metric
Macro-F1
Micro-F1
Time
Mem
Macro-F1
Micro-F1
Time
Mem
Macro-F1
Micro-F1
Time
Mem
Macro-F1
Micro-F1
Time
Mem
Co-GCN
31.5 ± 7.7
31.5 ± 7.7
5.0
128.0
85.5 ± 0.2
88.2 ± 0.2
37.7
201.8
97.2 ± 0.4
97.2 ± 0.4
7.4
216.0
32.0 ± 0.4
32.8 ± 0.4
4.5
196.3
PDMF
94.1 ± 0.7
94.2 ± 0.7
2.3
2834.2
55.8 ± 0.3
60.6 ± 0.2
11.2
3544.9
90.9 ± 0.0
90.9 ± 0.0
1.8
2973.6
47.0 ± 0.4
46.4 ± 0.4
11.1
2949.1
LGCNFF
OOM
OOM
-
-
OOM
OOM
-
-
OOM
OOM
-
-
OOM
OOM
-
-
ECMGD
OOM
OOM
-
-
OOM
OOM
-
-
OOM
OOM
-
-
OOM
OOM
-
-
Table 4. Classification results (mean% ± std%), inference time (s), and memory usage (MB) on large-scale datasets. The best and second-best results are highlighted in red and blue, respectively. “OOM” indicates out-of-memory errors.
Figure 2. t-SNE visualizations of embeddings learned by different methods on the heterogeneous graph dataset DBLP (top row), the multi-omics dataset TCGA (middle row), and the multi-view dataset HW (bottom row). A three-by-three grid of t-SNE scatter plots comparing learned embeddings. The rows correspond to DBLP, TCGA, and HW, respectively, while the columns compare two baseline methods with SynMDiff. Points belonging to different classes are distinguished by color. On DBLP, SynMDiff produces more compact and clearly separated class clusters than MHGCN and AMOGCN. On TCGA, the baseline embeddings exhibit substantial overlap, whereas SynMDiff separates most classes into distinct clusters. On HW, all methods form relatively clear clusters, with SynMDiff achieving slightly better separation. SynMDiff obtains the highest homogeneity and completeness scores in all three rows: 0.8154 and 0.8168 on DBLP, 0.7343 and 0.7467 on TCGA, and 0.9368 and 0.9374 on HW.
Figure 3. Ablation study of SynMDiff on six datasets. Six bar charts arranged in three rows and two columns report Micro-F1 scores on MITIndoor, HW, BRCA, UCEC, FreeBase, and Yelp. Each chart compares SynMDiff with three ablated variants that remove inter-view fusion, intra-view fusion, or the sampling mechanism. Error bars indicate variation across experimental runs. The complete SynMDiff model achieves the highest Micro-F1 score on every dataset. Removing any component reduces performance, with particularly large reductions on UCEC and FreeBase, demonstrating that inter-view fusion, intra-view fusion, and sampling all contribute to the final results.
Appendix figures & tables7 assets
Supplementary material from the paper’s appendix.
Appendix
Datasets
# Samples
#Features
#Subtypes
BRCA
511
mRNA: 1,000 CNV: 1,000 RPPA: 223
4
LGG
524
DNA: 2,000 mRNA: 2,000 miRNA: 548
2
UCEC
430
DNA: 2,000 mRNA: 2,000 miRNA: 554
3
GBMLGG
511
DNA: 2,000 mRNA: 2,000 miRNA: 548
3
TCGA
9,664
gene expression: 17,944 CNV: 17,944
28
Appendix
Table 5. A brief description of cancer subtype datasets.
Datasets
# Samples
#Features
#Views
Meta-Paths
#Classes
FreeBase
43,854
3,492
3
MAM MDM MWM
4
DBLP
27,194
334
3
APA APCPA APTPA
4
IMDB
12,722
1,232
3
MAM MDM MYM
3
Yelp
3,913
2,614
3
BUB BSB BLB
3
AMiner
55,783
128
2
PAP PRP
3
Appendix
Table 6. A brief description of heterogeneous graph datasets.
Datasets
# Samples
# Views
# Features
# Classes
HW
2,000
6
153/596/301/481/157/27
10
Youtube
2,000
6
2,000/1,024/64/512/64/647
10
Scene15
4,485
3
1,800/1,180/1,240
15
MITIndoor
5,360
4
3,600/1,770/1,240/4,096
67
IAPR
7,855
2
100/100
6
Caltech
9,144
6
48/40/254/1,984/512/928
102
Appendix
Table 7. A brief description of multi-view test datasets.
Dataset
learning rate
hops
k
BRCA
0.0005
3
10
LGG
0.0005
0
5
UCEC
0.0005
3
10
GBMLGG
0.0005
3
10
TCGA
0.0005
3
10
FreeBase
0.0005
3
10
Appendix
Table 8. Hyperparameter Configurations
Metric
Async.
w/o k -NN
w/o hi
SynMDiff
ACC (%)
96.12
95.81
96.44
97.06
Time (ms)
376.25
371.15
69.60
55.50
Mem. (MB)
5376.90
5666.89
99.99
101.27
Appendix
Table 9. Component-wise ablation results.
k
5
10
15
20
25
30
ACC (%)
95.81
97.13
97.25
96.25
96.31
95.88
Time (ms)
25.76
27.39
28.64
29.34
29.48
30.09
Mem. (MB)
99.52
100.71
99.61
100.73
99.54
99.54
Appendix
Table 10. Sensitivity to the neighborhood size k .
Figure 4. A WL-inspired comparison between asynchronous and synchronous multi-view fusion. Under asynchronous fusion, the missing edge between A2 and B2 interrupts the intermediate propagation pathway and isolates B2 from the complementary information in A1 . Synchronous fusion directly models the interaction between A1 and B2 , enabling the involved nodes to obtain the joint signature {1,2} . A toy graph compares asynchronous and synchronous fusion. A missing edge between A2 and B2 blocks indirect information flow from A1 to B2 in the asynchronous case. A direct cross-view interaction in the synchronous case allows the joint feature signature to reach all involved nodes.
Multi-view diffusion models have shown strong performance in scenes with strong geometric priors and sparse semantics, such as indoor rooms or simple outdoor environments (e.g., fields, courtyards). However, they often fail to maintain cross-view consistency under camera rotation, especially in structurally complex urban environments. Without explicit modeling of spherical correspondence across views, existing approaches tend to produce object duplication, structural distortion, and layout inconsistency. To address this limitation, we propose StreetDiff, a multi-view diffusion framework that explicitly enforces cross-view alignment during denoising. StreetDiff introduces a Panorama--Perspective Synergy design to decouple global layout reasoning from local detail synthesis, and incorporates a Panorama Alignment Module (PAM) that establishes spherical-projection-based attention constraints across views. By injecting structured alignment constraints without modifying the diffusion backbone, our framework achieves robust cross-view coherence in challenging urban street scene generation tasks. In addition, we construct Street360, a large-scale HDR multi-view urban panorama dataset. Extensive experiments demonstrate that StreetDiff significantly improves structural consistency and visual fidelity compared to prior multi-view diffusion generation methods.
Qi Zhang, Yanyifan Wang, Weiyuan Zhang +1
College of Computer Science and Software Engineering, Shenzhen University, China
Recent success in contrastive learning has sparked growing interest in more effectively leveraging multiple augmented views of data. While prior methods incorporate multiple views at the loss or feature level, they primarily capture pairwise relationships and fail to model the joint structure across all views. In this work, we propose a divergence-based similarity function (DSF) that explicitly captures the joint structure by representing each set of augmented views as a distribution and measuring similarity as the divergence between distributions. Extensive experiments demonstrate that DSF consistently improves performance across diverse tasks, including kNN classification, linear evaluation, transfer learning, and distribution shift, while also achieving greater efficiency than other multi-view methods. Furthermore, we establish a connection between DSF and cosine similarity, and demonstrate that, unlike cosine similarity, DSF operates effectively without the need for tuning a temperature hyperparameter.
Jaehyoung Jeon, Cheolsu Lim, Myungjoo Kang
Department of Mathematics, Seoul National University, Seoul, Korea · Research Institute of Mathematics, Seoul National University, Seoul, Korea
Recent advances in video diffusion models have enabled either long single-view generation through temporal autoregression, or short multi-view synthesis through bidirectional attention. However, generating long, multi-view consistent videos of dynamic scenes remains unsolved. In this work, we present MV-Forcing, a framework that composes temporal and view-wise autoregression within a single diffusion model by introducing a 4D geometric bridge between sequentially generated views. Our key insight is that an autoregressive 3D reconstruction model naturally interfaces between autoregressively generated views. Given a completed source view, we reconstruct its 3D structure and render a geometric prior of the next target viewpoint, which the diffusion model refines into a high-quality video. To extend generation beyond the teacher's fixed temporal window, we introduce a joint denoising regime where both view slots are initialized from noise during training, enabling temporally unbounded generation. We distill the model via Distribution Matching Distillation with Spatio-Temporal Self-Forcing, closing the train-inference exposure bias gap for both temporal and view-sequential autoregression. Extensive experiments on both synthetic and real-world data demonstrate that MV-Forcing produces geometrically consistent multi-view videos of dynamic scenes at arbitrary lengths and viewpoint counts using a single few-step student model.
Gal Fiebelman, Hadar Averbuch-Elor, Sagie Benaim
The Hebrew University of Jerusalem · Cornell University