Organizations: IROOTECH TECHNOLOGY, Hangzhou, Zhejiang, China · Wolf 1069 b Lab, Sany Group, Hangzhou, Zhejiang, China · Zhejiang University, Hangzhou, Zhejiang, China · IROOTECH TECHNOLOGY, Guangzhou, Guangdong, China · Wolf 1069 b Lab, Sany Group, Guangzhou, Guangdong, China · Central South University, Changsha, Hunan, China · BPIT, Sany Group, Changsha, Hunan, China
We propose FusionBERT, a novel multi-view visual fusion framework for image--3D multimodal retrieval. Existing image--3D representation learning methods predominantly focus on feature alignment of a single object image and its 3D model, limiting their applicability in realistic scenarios where an object is typically observed and captured from multiple viewpoints. Although multi-view observations naturally provide complementary geometric and appearance cues, existing multimodal large models rarely explore how to effectively fuse such multi-view visual information for better cross-modal retrieval. To address this limitation, we introduce a multi-view image--3D retrieval framework named FusionBERT, which innovatively utilizes a cross-attention-based multi-view visual aggregator to adaptively integrate features from multi-view images of an object. The proposed multi-view visual encoder fuses inter-view complementary relationships and selectively emphasizes informative visual cues across multiple views to get a more robustly fused visual feature for better 3D model matching. Furthermore, FusionBERT proposes a normal-aware 3D model encoder that can further enhance the 3D geometric feature of an object model by jointly encoding point normals and 3D positions, enabling a more robust representation learning for textureless or color-degraded 3D models. Extensive image--3D retrieval experiments on both synthetic 3D models and real-world industrial mechanical objects demonstrate that FusionBERT achieves significantly higher retrieval accuracy than SOTA multimodal large models under both single-view and multi-view settings, establishing a strong baseline for multi-view multimodal retrieval.
Figures & tables
Figure 1: An example of utilizing our FusionBERT model in images-3D model retrieval task with multi-view images as input query. Our FusionBERT model achieves a successful Top-1 retrieval, surpassing other SOTA image–3D retrieval models such as TAMM ( Zhang et al. 2024 ) , OpenShape ( Liu et al. 2023 ) , ULIP-2 ( Xue et al. 2024 ) , ReCon ( Qi et al. 2023 ) and Uni3D ( Zhou et al. 2024 ) in matching multi-view images to the right 3D model.
Figure 2: System overview. FusionBERT fine-tunes a multi-view fusion aggregator and a normal-aware 3D encoder by aligning fused multi-view features with 3D representations via contrastive learning, mitigating the cross-modal gap between images and 3D. Following TAMM ( Zhang et al. 2024 ) , we introduce IAA and TAA into the 3D encoder to decouple representations into visual and semantic subspaces, enabling more effective tri-modal pre-training. Contrastive learning maximizes similarity between matched features while minimizing those for mismatched ones.
Figure 3: Architecture of the normal-aware 3D model encoder. The input point cloud ( N×9 ) is partitioned into P patches via FPS and k NN grouping. Each patch is encoded by a Mini-PointNet into patch tokens, which are concatenated with a [CLS] token and processed by Transformer blocks. The output [CLS] token is projected to the global shape feature.
Figure 4: Two exemplar retrieval tasks on Objaverse-LVIS dataset ( Deitke et al. 2023 ) , where our FusionBERT model achieves the best performance with the correct result at Recall@1 for both cases, outperforming SOTA TAMM ( Zhang et al. 2024 ) , OpenShape ( Liu et al. 2023 ) , ULIP-2 ( Xue et al. 2024 ) , ReCon ( Qi et al. 2023 ) and Uni3D ( Zhou et al. 2024 ) .
Views
Pre-Trained Model
Objaverse-LVIS
LVIS no-RGB
ModelNet40
IMP
Top-1
Top-3
Top-5
Top-1
Top-3
Top-5
Top-1
Top-3
Top-5
Top-1
Top-3
Top-5
1 View
FusionBERT
52.10
70.36
76.96
26.78
43.56
52.08
8.55
17.85
23.39
19.25
41.71
53.74
TAMM
48.71
67.26
74.59
22.65
38.44
46.65
7.63
16.83
22.42
19.25
39.04
52.14
ReCon
45.19
64.06
72.37
21.73
37.35
45.70
8.01
17.82
22.81
18.71
39.04
51.34
OpenShape
43.37
62.56
70.47
19.18
33.87
41.95
6.18
15.27
20.70
11.76
34.22
44.12
ULIP-2
42.51
59.97
67.02
26.47
42.37
50.43
7.20
13.98
18.82
14.71
30.48
41.44
Table 1: Overall comparative experiments of our FusionBERT with SOTA pre-trained models OpenShape ( Liu et al. 2023 ) , TAMM ( Zhang et al. 2024 ) , ULIP-2 ( Xue et al. 2024 ) , ReCon ( Qi et al. 2023 ) and Uni3D ( Zhou et al. 2024 ) for image–3D retrieval task under both single-view and multi-view configurations. Evaluation is conducted on four datasets: Objaverse-LVIS ( Deitke et al. 2023 ) , Objaverse-LVIS without RGB (LVIS no-RGB), ModelNet40 ( Wu et al. 2015 ) and our self-captured Industrial Machinery Part (IMP) dataset. All values represent retrieval accuracy ( % ).
Modules
Objaverse-LVIS
MVVA
NA3DE
Top-1
Top-3
Top-5
Top-10
60.90
79.16
85.12
91.35
✓
65.74
82.29
87.64
92.73
✓
63.15
79.87
85.55
91.19
✓
✓
68.73
84.46
89.16
93.79
Table 2: Ablation study on multi-view visual aggregator (MVVA) and normal-aware 3D encoder (NA3DE) modules, with ✓marking enabled modules. We report retrieval accuracy ( % ) on Objaverse-LVIS with 3-view inputs. The multi-view aggregator defaults to mean pooling when disabled.
Fusion Method
Objaverse-LVIS
Top-1
Top-3
Top-5
Top-10
Mean Pooling
63.00
79.79
85.4
91.14
Max Pooling
49.97
63.96
68.53
73.33
Min Pooling
58.73
76.49
82.67
89.24
Weighted Pooling
64.35
81.08
86.87
92.65
Transformer Fusion
68.49
83.91
88.76
93.31
Table 3: Ablation study of different multi-view feature fusion methods. We report retrieval accuracy ( % ) on the Objaverse-LVIS dataset using 3-view inputs.
Figure 5: Ablation study on the number of input views for FusionBERT, OpenShape ( Liu et al. 2023 ) , TAMM ( Zhang et al. 2024 ) , ULIP-2 ( Xue et al. 2024 ) , ReCon ( Qi et al. 2023 ) and Uni3D ( Zhou et al. 2024 ) across 1∼12 views with Top-5 retrieval accuracies ( % ) on Objaverse-LVIS.
School of Engineering Mathematics and Technology, University of Bristol, Bristol, UK · Department of Automation, Tsinghua University, Beijing, China · UCAS-Terminus AI Lab, University of Chinese Academy of Sciences, Beijing 100049, China