Unsupervised domain adaptation (UDA) effectively bridges the domain gap between a labeled source domain and an unlabeled target domain, but assumes that the two domains share the same modality. Heterogeneous domain adaptation (HDA) instead handles different feature spaces across domains, yet requires labeled target samples or paired data linking the source and target domains. Neither applies when a labeled source domain and a fully unlabeled target domain each hold an entirely distinct modality (e.g., 2D images and 3D point clouds). To address this limitation, we introduce a new setting termed Heterogeneous-Modal Unsupervised Domain Adaptation (HMUDA), which transfers knowledge across modalities via an unlabeled bridge domain containing paired observations from both modalities, whose distribution may deviate from those of the source and target domains. To learn under the HMUDA setting, we propose Latent Space Bridging (LSB), a dual-branch framework where a feature consistency loss on paired bridge samples closes the modality gap and a class-centroid alignment loss reduces the source-target discrepancy. Extensive experiments on eight benchmark settings covering both 2D-to-3D and 3D-to-2D transfer demonstrate that LSB achieves state-of-the-art performance.
Figures & tables
Fig. 1: An illustration of the HMUDA setting.
Source
Target
Bridge
E-to-E
M1
M2
M1
M2
UDA
✓
✗
✓
✗
✗
✓
MM-UDA
✓
✓
✓
✓
✗
✓
HDA
✓
✗
✗
✓
✗
✗
HMUDA
✓
✗
✗
✓
✓
✓
TABLE I: Comparison between HMUDA and existing DA settings. M1 and M2 denote two modalities. ‘Bridge’ indicates whether a third unlabeled domain with paired modalities is available, and ‘E-to-E’ indicates whether the specific model is trained end-to-end from raw inputs.
Fig. 2: An illustration of the proposed LSB framework. Lines within different colors denote the data flow for computing different losses.
B
S
T
Scenario
Train
Train
Train
Val/Test
USA → Sing.
Sem.
18,029
15,695
9,665
2,770/2,929
Day → Night
24,745
2,779
602/602
Virt. → A2D2
2,126
24,461
808/2,426
USA → Sing.
A2D2
24,461
15,695
9,665
2,770/2,929
Day → Night
24,745
2,779
602/602
TABLE II: Statistics of the number of samples in each split of datasets for all eight settings. Each setting is evaluated in both the 2D-to-3D and 3D-to-2D directions, which share the same data splits and differ only in which modality serves as the source domain. Note that the training samples in the target and bridge domains are without labels.
USA → Sing.
Day → Night
Virt. → A2D2
USA → Sing.
Day → Night
Virt. → Sem.
Sem. → A2D2
A2D2 → Sem.
Avg.
Bridge Domain B
Sem.
A2D2
Virt.
–
2D → 3D
Oracle
77.69
73.71
71.50
77.69
73.71
82.75
71.50
82.75
76.41
xMUDA [ 65 ]
65.12
75.11
61.03
65.12
75.11
57.97
44.95
65.07
63.69
Source-Only
51.04
57.32
19.00
49.21
55.96
36.44
43.75
43.28
44.50
PL [ 5 ]
52.08
59.33
16.84
55.39
61.95
44.49
39.66
38.25
46.00
CDSPP [ 32 ]
17.61
21.81
9.74
17.61
21.81
12.44
11.27
11.00
15.41
TABLE III: Testing results on HMUDA tasks in terms of mIoU. Each domain transfer is evaluated in both 2D-to-3D and 3D-to-2D modality directions. The last column reports the average over all eight settings for each transfer direction. The best performance is in bold .
Lsegs
Lsegb
Lconb
Lali
USA → Sing. ( Sem. )
Day → Night ( Sem. )
Virt. → A2D2 ( Sem. )
Sem. → A2D2 ( Virt. )
✓
✓
✗
✗
53.62
61.86
22.72
42.47
✓
✓
✓
✗
52.68
61.12
28.24
45.27
✓
✓
✗
✓
54.34
60.86
27.96
46.37
✓
✓
✓
✓
56.41
62.25
33.37
46.34
Lsegs
Lsegb
Lconb
Lali
USA → Sing. ( A2D2 )
Day → Night ( A2D2 )
Virt. → Sem. ( A2D2 )
A2D2 → Sem. ( Virt. )
✓
✓
✗
✗
49.35
60.64
35.74
42.29
TABLE IV: Effect of losses Lsegs , Lsegb , Lconb , and Lali in terms of mIoU for 2D-to-3D HMUDA tasks. The bridge domain B for each task is given in parentheses. The best performance is in bold .
Fig. 3: Sensitivity analysis of the LSB method with respect to hyperparameters and the number of samples in the bridge domain.
USA
Day
A2D2
Method
B
→
→
B
→
Sing.
Night
Sem.
LSB (w/o pϕ )
A2D2
20.78
3.82
Virt.
15.00
LSB (w/o ph )
40.47
43.10
32.52
LSB
57.13
63.15
47.22
TABLE V: Effect of pϕ and ph on three 2D-to-3D HMUDA tasks.
USA
Day
A2D2
Method
B
→
→
B
→
Sing.
Night
Sem.
LSB ( w. Lali(B,T) )
A2D2
50.86
60.12
Virt.
44.44
LSB
57.13
63.15
47.22
TABLE VI: Effect of cross-modal alignment on three 2D-to-3D HMUDA tasks.
Fig. 4: Qualitative results on three 2D-to-3D HMUDA tasks: USA → Sing. , Day → Night , and A2D2 → Sem. . (⋅) in the vertical axis denotes the bridge domain B used in the HMUDA task. For example, USA → Sing.(Sem.) denotes the transfer from USA to Sing. via the bridge domain Sem. .
Fig. 5: t-SNE visualization of target feature embeddings for the Source-Only and LSB methods.