cs.CVApr 2, 2026

RGB-Pointmap Pretraining for Unified 3D Scene Understanding

Authors: Ye MaoWeixun LuoRanran HuangJunpeng JingKrystian Mikolajczyk

Organizations: Imperial College London

Abstract

Pretraining 3D encoders through alignment with Contrastive Language-Image Pre-training (CLIP) has emerged as a promising direction for learning generalizable representations for 3D scene understanding. In this paper, we propose UniScene3D, a transformer-based framework that learns unified 3D scene representations from multi-view RGB-Pointmap inputs by leveraging the priors of a pretrained 2D foundation model. For robust RGB-Pointmap representation learning, we introduce cross-view geometric alignment and grounded view alignment to enforce geometric and semantic consistency across views. Extensive low-shot and task-specific fine-tuning on viewpoint grounding, scene retrieval, scene classification, and 3D visual question answering demonstrates state-of-the-art performance. These results establish UniScene3D as an effective framework for unified 3D scene understanding. Project page: https://yebulabula.github.io/UniScene3D/

Explore similar work

CardsList
  1. Utonia: Toward One Encoder for All Point Clouds

    Mar 3, 2026Yujia Zhang, Xiaoyang Wu, Yunhan Yang +6Spatial ReasoningEncoders

  2. Multi-View Foundation Models

    Dec 17, 2025Leo Segre, Or Hirschorn, Shai Avidan3D Foundation ModelsMulti-View

  3. From Alignment to Fusion in 3D Vision-Language

    Sep 23, 2026Xueqi Qiu, Xingyu Miao, Jingjing Deng +3Vision-Language Alignment