cs.CVOct 1, 2026

PAGER: Partial-to-global Alignment via Geometric and Relational Distillation

Authors: Akira-Miranda Adeyomi Adeniran-Lowe, Binod Singh, Lars Arnold Dethlefsen, Lazaros Nalpantidis, Theodora Kontogianni

Organizations: Technical University of Denmark · Pioneer Center for Artificial Intelligence

Abstract

Pretrained 3D encoders are typically developed on globally reconstructed scenes expressed in a consistent world coordinate frame, whereas embodied systems must reason from partial, viewpoint-dependent observations in camera coordinates. We show that this shift from globally learned 3D feature spaces to realistic partial observations exposes a severe representation mismatch, which we find consistently across representative state-of-the-art encoders, including Sonata and Concerto. A frozen Sonata encoder with a global linear probe achieves 72.47 mIoU on full ScanNet scenes, but 2.57 mIoU on single-frame camera-coordinate inputs. Training-free gravity alignment recovers performance to 41.64 mIoU, showing that coordinate-frame mismatch is a dominant source of degradation but cannot be fully resolved through canonicalization alone. We introduce PAGER, a label-free adaptation method that aligns partial-view features with a frozen global 3D semantic space using only paired partial/global geometry. It learns lightweight adaptation modules while keeping the pretrained encoder and global segmentation probe frozen. Matched-point feature alignment anchors partial features to their global counterparts, while relational supervision preserves their similarity structure with respect to the global representation. Global geometry provides supervision only during training. Inference operates directly on the partial observation. Without partial-view labels, PAGER outperforms label-supervised PEFT on both Sonata and Concerto, and in zero-shot ScanNet→\rightarrowScanNet++ transfer surpasses fully fine-tuned Sonata (53.9353.93 vs.\ 48.0948.09 mIoU), suggesting that preserving the frozen global representation can improve cross-dataset transfer.

Figures & tables

Appendix figures & tables12 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Robust 3D Alignment of Generative Reconstructions via Partial Monocular Observations

    Jul 1, 2026Yuchen Zhang, Luanyuan Dai, Yiwei Wang +73D ReconstructionMonocular

  2. Less Context, Better Geometry: Masked Geometric Encoder for Robust 3D Foundation Models

    Oct 5, 2026Zhimin Shao, Xijun Liu, Zhaoliang Zhang +43D Foundation Models