cs.CVOct 5, 2026

LoDEOT: Low-Dimensional and Efficient Offset Tokens for Building Footprint Extraction from Off-Nadir Imagery

Authors: Kai Li, Zigan Zhou, Zhenyang Li, Hui Shan, Zhe Chen, Yupeng Deng, Zhihao Xi, Yu Meng, +2 more

Organizations: City University of Hong Kong · University of the Chinese Academy of Sciences · University of Hong Kong · Zhejiang University · University of Southampton

Abstract

Instance-level roof-to-footprint offset (RFO) prediction is central to extracting building footprints from off-nadir imagery. Query-based pipelines commonly use high-dimensional instance tokens to predict signed two-dimensional RFOs. We investigate whether RFO prediction can instead use a compact offset token. Under local pinhole projection and vertical-extrusion assumptions, the idealized RFO map admits a five-parameter sufficient descriptor comprising intrinsic shape, composite amplitude, and relative geometry. This factorization provides a structural prior for a five-dimensional offset token, whose channels learn task-relevant latent representations through end-to-end training. Based on this design, we propose LoDEOT, which retains high-dimensional instance tokens for detection and segmentation but maps instance-token, concentration-gated roof, and box-mask evidence to a five-dimensional offset token followed by an independent two-dimensional readout. Known denoising-query target indices further align each supervised decoder-layer estimate with the same clean instance RFO, organizing successive predictions as target-aligned recovery under perturbed query conditions. Experiments on five real-world building datasets demonstrate the effectiveness of LoDEOT for building footprint extraction. Experiments on real-world building datasets demonstrate that a five-dimensional offset token can support accurate RFO prediction. On BONAI, LoDEOT achieves the best roof-detection bAP and bAP50 and leads all five offset-corrected footprint metrics among the evaluated end-to-end methods, with FAP50 of 54.58 and mEPE of 5.23 pixels. Its FAP50 exceeds those of the evaluated end-to-end baselines by 7.56-16.85 percentage points.

Figures & tables

Appendix figures & tables1 asset

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Joint Instance Segmentation and Geometric Attribute Regression for Roof Structures in Aerial Imagery

    May 25, 2026Luuk Versteeg, Rob G. J. Wijnhoven, Martin R. OswaldAerial ImageryLarge-Scale 3D Editing Dataset

  2. PCFootprint: A Large-Scale Dataset and Benchmark for Vectorized Building Footprint Extraction from Aerial LiDAR Point Clouds

    Jun 18, 2026Haoyuan Shen, Kuihao Wang, Ruisheng Wang +1FootprintPoint Clouds

  3. ObliCity: A Benchmark and Baseline for Roof-to-Ground Projection Displacement Correction

    Jul 28, 2026Kai Li, Yupeng Deng, Ligao Deng +6ContoursUrban Environments