Organizations: School of Computer Science, National Engineering Research Center for Multimedia Software, Institute of Artificial Intelligence, Hubei Key Laboratory of Multimedia and Network Communication Engineering, Wuhan University · University of North Texas · Institute of Software, Chinese Academy of Sciences
3D Visual Query Localization (3DVQL) retrieves the latest contiguous occurrence of a queried object in an RGB--point-cloud sequence and predicts a 9-DoF cuboid for every response frame. The query is captured independently of the search sequence, so its annotated pose may differ from how the object appears in the search frames. The benchmark baseline predicts cuboids after feature modeling, leaving their geometry unused for subsequent feature refinement. We investigate whether complete intermediate cuboids can improve query and proposal representations before final decoding. We introduce Progressive Geometric Learning for 3DVQL (PGL-3D), a predict--select--refine--re-predict framework that uses intermediate cuboids to guide the aggregation of search evidence and update query and proposal representations. A shared head first predicts a complete cuboid for every proposal. Query--Tube--Memory (QTM) then selects reference observations by combining proposal association, cuboid quality, frame response, and target absence, since association confidence alone establishes neither target presence nor geometric accuracy. The center, size, and orientation of each selected cuboid define soft pooling weights over query-conditioned proposal features. The pooled memory updates the query and proposal representations, and the head re-predicts from the updated features. A training-only objective, ST-D9O, supervises cuboid geometry at every stage by adding boundary, signed-distance, and soft-overlap terms to parameter regression. PGL-3D achieves a mean stAP of 0.270±0.004 on 3DVQL, compared with 0.044 reported for LaF. Ablations support the benefits of geometry-guided feature updates, while stage-wise analyses show improved cuboid accuracy. Replacing the geometry objective in our PROT3D reproduction with ST-D9O improves mAO on GSOT3D from 21.63% to 25.78%. Our code and models will be released.
Visual query localization (VQL) aims to predict the spatio-temporal response of the most recent occurrence in a sequence given a query. Currently, most research focuses on visual query localization in 2D videos, while its counterpart in 3D space has received little attention. In this paper, we make the first attempt to address visual query localization in the 3D world by introducing a novel benchmark, dubbed 3DVQL. Specifically, 3DVQL contains 2,002 sequences with around 170,000 frames and 6.4K response track segments from 38 object categories. Each sequence in 3DVQL is provided with multiple modalities, including point clouds, RGB images, and depth images, to support flexible research. To ensure high-quality annotations, each sequence is manually annotated with multiple rounds of verification and refinement. To the best of our knowledge, 3DVQL is the first benchmark for 3D multimodal visual query localization. To facilitate comparison in subsequent research, we implement a series of representative 3D multimodal VQL baselines using point clouds and RGB images. The experimental results show that existing methods exhibit significant performance variations across different fusion modules. To encourage future research, we propose a lift-and-attention fusion algorithm named LaF, which significantly outperforms existing baseline models. Our benchmark and model will be publicly released at https://github.com/wuhengliangliang/3DVQL.
Liang Peng, Bohan Tan, Zhipeng Zhang +4
Wuhan University · AutoLab, SAI, Shanghai Jiao Tong University · Anyverse Dynamics +2
Visual query localization (VQL) aims to retrieve and re-localize a queried object in egocentric videos, yet remains challenging when object boundaries are ambiguous and global context cannot effectively guide fine-grained localization. Human vision handles such ambiguity through a hierarchical process: it rapidly screens foreground candidates, selectively attends to the target despite distractors, refines perception via feedback between global context and local detail, and, when a single view is unreliable, integrates evidence across viewpoints according to its credibility. Inspired by these competencies, we propose \textbf{EgoHieraLoc}, a unified framework for VQL-2D and VQL-3D. A Discriminative Parsing Module first extracts foreground-aware query representations using segmentation priors; a Query-Aware Module then performs robust target localization through discriminative correlation filtering with deformable modeling; and a Regional Adaptation Module feeds multi-scale context back into local regions to recover precise object boundaries. To extend this perceptual hierarchy to 3D localization, we introduce Geometric-Semantic Joint Confidence (GSJC), which multiplicatively couples segmentation confidence with local depth consistency, multi-view back-projection consistency, and triangulation-baseline quality, so that a viewpoint contributes to the 3D estimate only when it is credible both semantically and geometrically. Extensive experiments demonstrate state-of-the-art performance on both VQL-2D and -3D benchmarks.
Yifei Cao, Guolong Wang, Mingliang Hou +3
Dalian University of Technology, Dalian, China · University of International Business and Economics, Beijing, China · Jinan University, Guangzhou, China
Language-based 3D localization retrieves the point-cloud submap containing a target position from descriptions of nearby objects and their spatial relations. Existing methods typically compress queries and submaps into global descriptors, potentially obscuring object-level semantics and cross-description spatial coherence. We propose Position-Conditioned Evidence Localization (PosEviLoc), a query-position-aware framework for coarse text-to-point-cloud localization. Instead of relying on global matching, PosEviLoc evaluates each candidate submap using explicit semantic and spatial evidence. It models direction as a relation jointly determined by an object position and a hypothetical query position. The resulting Query-Position Spatial Evidence Field (QSEF) measures the fraction of query descriptions supported at each hypothetical position, explicitly capturing their agreement without using the ground-truth query pose to construct the evidence field. A Multi-Level Evidence Readout (MER) summarizes this evidence in a compact representation, which a lightweight MLP converts into a retrieval score. Across five benchmarks, PosEviLoc outperforms MNCL by an average of 17 percentage points in Recall@1. When used as a plug-and-play reranker, it improves MNCL by an average of 16 percentage points. Moreover, PosEviLoc introduces substantially fewer parameters and achieves faster inference speed than existing methods.
Tianyi Shang, Yike Shi, Zhenyu Li
Purdue University · Xi’an Jiaotong-Liverpool University · Qilu University of Technology (Shandong Academy of Sciences)