Organizations: IROOTECH TECHNOLOGY, Hangzhou, Zhejiang, China · Wolf 1069 b Lab, Sany Group, Hangzhou, Zhejiang, China · Zhejiang University, Hangzhou, Zhejiang, China · IROOTECH TECHNOLOGY, Guangzhou, Guangdong, China · Wolf 1069 b Lab, Sany Group, Guangzhou, Guangdong, China · Central South University, Changsha, Hunan, China · BPIT, Sany Group, Changsha, Hunan, China
We propose FusionBERT, a novel multi-view visual fusion framework for image--3D multimodal retrieval. Existing image--3D representation learning methods predominantly focus on feature alignment of a single object image and its 3D model, limiting their applicability in realistic scenarios where an object is typically observed and captured from multiple viewpoints. Although multi-view observations naturally provide complementary geometric and appearance cues, existing multimodal large models rarely explore how to effectively fuse such multi-view visual information for better cross-modal retrieval. To address this limitation, we introduce a multi-view image--3D retrieval framework named FusionBERT, which innovatively utilizes a cross-attention-based multi-view visual aggregator to adaptively integrate features from multi-view images of an object. The proposed multi-view visual encoder fuses inter-view complementary relationships and selectively emphasizes informative visual cues across multiple views to get a more robustly fused visual feature for better 3D model matching. Furthermore, FusionBERT proposes a normal-aware 3D model encoder that can further enhance the 3D geometric feature of an object model by jointly encoding point normals and 3D positions, enabling a more robust representation learning for textureless or color-degraded 3D models. Extensive image--3D retrieval experiments on both synthetic 3D models and real-world industrial mechanical objects demonstrate that FusionBERT achieves significantly higher retrieval accuracy than SOTA multimodal large models under both single-view and multi-view settings, establishing a strong baseline for multi-view multimodal retrieval.
Figures & tables
Figure 1: An example of utilizing our FusionBERT model in images-3D model retrieval task with multi-view images as input query. Our FusionBERT model achieves a successful Top-1 retrieval, surpassing other SOTA image–3D retrieval models such as TAMM ( Zhang et al. 2024 ) , OpenShape ( Liu et al. 2023 ) , ULIP-2 ( Xue et al. 2024 ) , ReCon ( Qi et al. 2023 ) and Uni3D ( Zhou et al. 2024 ) in matching multi-view images to the right 3D model.
Figure 2: System overview. FusionBERT fine-tunes a multi-view fusion aggregator and a normal-aware 3D encoder by aligning fused multi-view features with 3D representations via contrastive learning, mitigating the cross-modal gap between images and 3D. Following TAMM ( Zhang et al. 2024 ) , we introduce IAA and TAA into the 3D encoder to decouple representations into visual and semantic subspaces, enabling more effective tri-modal pre-training. Contrastive learning maximizes similarity between matched features while minimizing those for mismatched ones.
Figure 3: Architecture of the normal-aware 3D model encoder. The input point cloud ( N×9 ) is partitioned into P patches via FPS and k NN grouping. Each patch is encoded by a Mini-PointNet into patch tokens, which are concatenated with a [CLS] token and processed by Transformer blocks. The output [CLS] token is projected to the global shape feature.
Figure 4: Two exemplar retrieval tasks on Objaverse-LVIS dataset ( Deitke et al. 2023 ) , where our FusionBERT model achieves the best performance with the correct result at Recall@1 for both cases, outperforming SOTA TAMM ( Zhang et al. 2024 ) , OpenShape ( Liu et al. 2023 ) , ULIP-2 ( Xue et al. 2024 ) , ReCon ( Qi et al. 2023 ) and Uni3D ( Zhou et al. 2024 ) .
Views
Pre-Trained Model
Objaverse-LVIS
LVIS no-RGB
ModelNet40
IMP
Top-1
Top-3
Top-5
Top-1
Top-3
Top-5
Top-1
Top-3
Top-5
Top-1
Top-3
Top-5
1 View
FusionBERT
52.10
70.36
76.96
26.78
43.56
52.08
8.55
17.85
23.39
19.25
41.71
53.74
TAMM
48.71
67.26
74.59
22.65
38.44
46.65
7.63
16.83
22.42
19.25
39.04
52.14
ReCon
45.19
64.06
72.37
21.73
37.35
45.70
8.01
17.82
22.81
18.71
39.04
51.34
OpenShape
43.37
62.56
70.47
19.18
33.87
41.95
6.18
15.27
20.70
11.76
34.22
44.12
ULIP-2
42.51
59.97
67.02
26.47
42.37
50.43
7.20
13.98
18.82
14.71
30.48
41.44
Table 1: Overall comparative experiments of our FusionBERT with SOTA pre-trained models OpenShape ( Liu et al. 2023 ) , TAMM ( Zhang et al. 2024 ) , ULIP-2 ( Xue et al. 2024 ) , ReCon ( Qi et al. 2023 ) and Uni3D ( Zhou et al. 2024 ) for image–3D retrieval task under both single-view and multi-view configurations. Evaluation is conducted on four datasets: Objaverse-LVIS ( Deitke et al. 2023 ) , Objaverse-LVIS without RGB (LVIS no-RGB), ModelNet40 ( Wu et al. 2015 ) and our self-captured Industrial Machinery Part (IMP) dataset. All values represent retrieval accuracy ( % ).
Modules
Objaverse-LVIS
MVVA
NA3DE
Top-1
Top-3
Top-5
Top-10
60.90
79.16
85.12
91.35
✓
65.74
82.29
87.64
92.73
✓
63.15
79.87
85.55
91.19
✓
✓
68.73
84.46
89.16
93.79
Table 2: Ablation study on multi-view visual aggregator (MVVA) and normal-aware 3D encoder (NA3DE) modules, with ✓marking enabled modules. We report retrieval accuracy ( % ) on Objaverse-LVIS with 3-view inputs. The multi-view aggregator defaults to mean pooling when disabled.
Fusion Method
Objaverse-LVIS
Top-1
Top-3
Top-5
Top-10
Mean Pooling
63.00
79.79
85.4
91.14
Max Pooling
49.97
63.96
68.53
73.33
Min Pooling
58.73
76.49
82.67
89.24
Weighted Pooling
64.35
81.08
86.87
92.65
Transformer Fusion
68.49
83.91
88.76
93.31
Table 3: Ablation study of different multi-view feature fusion methods. We report retrieval accuracy ( % ) on the Objaverse-LVIS dataset using 3-view inputs.
Figure 5: Ablation study on the number of input views for FusionBERT, OpenShape ( Liu et al. 2023 ) , TAMM ( Zhang et al. 2024 ) , ULIP-2 ( Xue et al. 2024 ) , ReCon ( Qi et al. 2023 ) and Uni3D ( Zhou et al. 2024 ) across 1∼12 views with Top-5 retrieval accuracies ( % ) on Objaverse-LVIS.
Foundation models are vital tools in various Computer Vision applications. They take as input a single RGB image and output a deep feature representation that is useful for various applications. However, in case we have multiple views of the same 3D scene, they operate on each image independently and do not always produce consistent features for the same 3D point. We propose a way to convert a Foundation Model into a Multi-View Foundation Model. Such a model takes as input a set of images and outputs a feature map for each image such that the features of corresponding points are as consistent as possible. This approach bypasses the need to build a consistent 3D model of the features and allows direct manipulation in the image space. Specifically, we show how to augment Transformers-based foundation models (i.e., DINO, SAM, CLIP) with intermediate 3D-aware attention layers that help match features across different views. As leading examples, we show surface normal estimation and multi-view segmentation tasks. Quantitative experiments show that our method improves feature matching considerably compared to current foundation models.
The bird's-eye view (BEV) representation enables multi-sensor features to be fused within a unified space, serving as the primary approach for achieving comprehensive 3D perception. However, the discrete grid representation of BEV leads to significant detail loss and limits feature alignment and cross-modal information interaction in multimodal fusion perception. In this work, we break from the conventional BEV paradigm and propose a new universal framework for multi-modal fusion based on 3D Gaussian representation. This approach naturally unifies multi-modal features within a shared and continuous 3D Gaussian space, effectively preserving edge and fine texture details. To achieve this, we design a novel forward-projection-based multi-modal Gaussian initialization module and a shared cross-modal Gaussian encoder that iteratively updates Gaussian properties based on an attention mechanism. GaussianFusion is inherently a task-agnostic model, with its unified Gaussian representation naturally supporting various 3D perception tasks. Extensive experiments demonstrate the generality and robustness of GaussianFusion. On the nuScenes dataset, it outperforms the 3D object detection baseline BEVFusion by 2.6 NDS. Its variant surpasses GaussFormer on 3D semantic occupancy with 1.55 mIoU improvement while using only 30% of the Gaussians and achieving a 450% speedup.
Unified 3D vision-language systems must combine complementary geometry, scale, and appearance cues while supporting tasks from instance segmentation to language-guided reasoning. Existing methods often process point clouds, voxel grids, and multi-view images independently; directly combining these heterogeneous representations may leave substantial feature discrepancy unresolved, while subsequent unconstrained adaptation may distort their internal geometry. We propose an align-then-fuse framework that first applies triple pairwise cosine alignment to establish segment-level correspondence across the three representations and then retrieves task-conditioned features with a prompt-guided query decoder. Before fusion, representation-specific query features are transformed by learnable mappings constrained to the special orthogonal group. These mappings preserve inner products and Euclidean distances within each representation, permitting controlled representation-specific re-parameterisation without arbitrarily distorting its internal geometry. The transformed features are subsequently combined through Adaptive Fusion under downstream task supervision. Experiments cover eight datasets for instance segmentation, visual grounding, question answering, and dense captioning. Compared with PQ3D, the model improves average precision by 3.2 points on ScanNet200 and grounding accuracy by 2.9, 10.6, 4.6, and 4.1 points on ScanRefer, Nr3D, Sr3D, and Multi3DRefer, respectively, while also improving performance on ScanQA, SQA3D, and Scan2Cap. Ablations further support the complementary roles of alignment and orthogonal re-parameterisation and the effectiveness of Adaptive Fusion.
Xueqi Qiu, Xingyu Miao, Jingjing Deng +3
School of Engineering Mathematics and Technology, University of Bristol, Bristol, UK · Department of Automation, Tsinghua University, Beijing, China · UCAS-Terminus AI Lab, University of Chinese Academy of Sciences, Beijing 100049, China