cs.CVSep 27, 2026

PGL-3D: Towards Progressive Geometric Learning for 3D Visual Query Localization

Authors: Liang Peng, Shizhuo Mu, Bohan Tan, Wenyuan Wang, Chen Zhao, Xingping Dong, Heng Fan, Libo Zhang, +1 more

Organizations: School of Computer Science, National Engineering Research Center for Multimedia Software, Institute of Artificial Intelligence, Hubei Key Laboratory of Multimedia and Network Communication Engineering, Wuhan University · University of North Texas · Institute of Software, Chinese Academy of Sciences

Abstract

3D Visual Query Localization (3DVQL) retrieves the latest contiguous occurrence of a queried object in an RGB--point-cloud sequence and predicts a 9-DoF cuboid for every response frame. The query is captured independently of the search sequence, so its annotated pose may differ from how the object appears in the search frames. The benchmark baseline predicts cuboids after feature modeling, leaving their geometry unused for subsequent feature refinement. We investigate whether complete intermediate cuboids can improve query and proposal representations before final decoding. We introduce Progressive Geometric Learning for 3DVQL (PGL-3D), a predict--select--refine--re-predict framework that uses intermediate cuboids to guide the aggregation of search evidence and update query and proposal representations. A shared head first predicts a complete cuboid for every proposal. Query--Tube--Memory (QTM) then selects reference observations by combining proposal association, cuboid quality, frame response, and target absence, since association confidence alone establishes neither target presence nor geometric accuracy. The center, size, and orientation of each selected cuboid define soft pooling weights over query-conditioned proposal features. The pooled memory updates the query and proposal representations, and the head re-predicts from the updated features. A training-only objective, ST-D9O, supervises cuboid geometry at every stage by adding boundary, signed-distance, and soft-overlap terms to parameter regression. PGL-3D achieves a mean stAP of 0.270±0.0040.270 \pm 0.004 on 3DVQL, compared with 0.0440.044 reported for LaF. Ablations support the benefits of geometry-guided feature updates, while stage-wise analyses show improved cuboid accuracy. Replacing the geometry objective in our PROT3D reproduction with ST-D9O improves mAO on GSOT3D from 21.63%21.63\% to 25.78%25.78\%. Our code and models will be released.

Explore similar work

CardsList
  1. Towards Visual Query Localization in the 3D World

    May 2, 2026Liang Peng, Bohan Tan, Zhipeng Zhang +4Sequential Visual Localization3D World

  2. EgoHieraLoc: A Cortically Inspired Hierarchical Segmentation-Guided Framework for Egocentric Visual Query Localization

    Aug 10, 2026Yifei Cao, Guolong Wang, Mingliang Hou +3Sequential Visual LocalizationIndoor Localization

  3. PosEviLoc: Position-Conditioned Spatial Evidence for Language-Based 3D Localization

    Sep 20, 2026Tianyi Shang, Yike Shi, Zhenyu LiIndoor LocalizationPoint Clouds