LoDEOT: Low-Dimensional and Efficient Offset Tokens for Building Footprint Extraction from Off-Nadir Imagery
Organizations: City University of Hong Kong · University of the Chinese Academy of Sciences · University of Hong Kong · Zhejiang University · University of Southampton
Abstract
Instance-level roof-to-footprint offset (RFO) prediction is central to extracting building footprints from off-nadir imagery. Query-based pipelines commonly use high-dimensional instance tokens to predict signed two-dimensional RFOs. We investigate whether RFO prediction can instead use a compact offset token. Under local pinhole projection and vertical-extrusion assumptions, the idealized RFO map admits a five-parameter sufficient descriptor comprising intrinsic shape, composite amplitude, and relative geometry. This factorization provides a structural prior for a five-dimensional offset token, whose channels learn task-relevant latent representations through end-to-end training. Based on this design, we propose LoDEOT, which retains high-dimensional instance tokens for detection and segmentation but maps instance-token, concentration-gated roof, and box-mask evidence to a five-dimensional offset token followed by an independent two-dimensional readout. Known denoising-query target indices further align each supervised decoder-layer estimate with the same clean instance RFO, organizing successive predictions as target-aligned recovery under perturbed query conditions. Experiments on five real-world building datasets demonstrate the effectiveness of LoDEOT for building footprint extraction. Experiments on real-world building datasets demonstrate that a five-dimensional offset token can support accurate RFO prediction. On BONAI, LoDEOT achieves the best roof-detection bAP and bAP50 and leads all five offset-corrected footprint metrics among the evaluated end-to-end methods, with FAP50 of 54.58 and mEPE of 5.23 pixels. Its FAP50 exceeds those of the evaluated end-to-end baselines by 7.56-16.85 percentage points.
Figures & tables
| Method | Roof Detection and Segmentation | Offset-Corrected Footprint | |||||||
|---|---|---|---|---|---|---|---|---|---|
| bAP | bAP50 | sAP | sAP50 | FAP50 | F1 | B-IoU | mEPE | PCP@5 | |
| Externally Guided Methods | |||||||||
| OBM | – | – | 12.59 | 36.71 | 22.21 | 41.72 | 9.25 | 5.72 | 53.68 |
| PolyFootNet | – | – | 8.91 | 33.18 | 18.76 | 39.75 | 7.32 | 6.54 | 44.71 |
| End-to-End Methods | |||||||||
| LOFT | 10.50 | 42.90 | 21.00 | 59.50 | 37.73 | 49.64 | 8.78 | 5.65 | 57.74 |
| Level | BONAI | IRSA | WHU | UAV | Huizhou | Average | ||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| sAP50 | FAP50 | sAP50 | FAP50 | sAP50 | FAP50 | sAP50 | FAP50 | sAP50 | FAP50 | sAP50 | FAP50 | |
| L1 | 54.94 | 54.58 | 9.70 | 8.77 | 49.50 | 48.75 | 31.64 | 28.49 | 35.72 | 31.17 | 36.30 | 34.35 |
| L2 | 54.68 | 51.58 | 51.67 | 48.76 | 26.65 | 26.60 | 14.03 | 12.98 | 8.82 | 5.56 | 31.17 | 29.09 |
| L3 | 54.21 | 51.76 | 49.42 | 46.30 | 96.08 | 96.08 | 24.25 | 22.65 | 31.22 | 25.01 | 51.04 | 48.36 |
| L4 | 55.22 | 52.29 | 49.27 | 46.34 | 95.55 | 95.54 | 44.99 | 41.75 | 39.13 | 32.28 | 56.83 | 53.64 |
| bAP | bAP50 | sAP | sAP50 | FAP50 | F1 | |
|---|---|---|---|---|---|---|
| 2 | ||||||
| 3 | ||||||
| 4 | ||||||
| 5 | ||||||
| 6 | ||||||
| 7 |
| Configuration | bAP | bAP50 | sAP | sAP50 | FAP50 | F1 | B-IoU | mEPE | PCP@5 |
|---|---|---|---|---|---|---|---|---|---|
| Disabled | 24.02 | 40.59 | 23.06 | 40.57 | 37.21 | 43.91 | 21.13 | 3.29 | 84.58 |
| Multiplicative | 28.33 | 47.32 | 27.49 | 47.66 | 44.21 | 50.24 | 22.29 | 3.04 | 85.52 |
| Additive | 30.27 | 49.57 | 29.17 | 49.65 | 46.79 | 52.54 | 22.27 | 2.71 | 88.09 |
Appendix figures & tables1 asset
Supplementary material from the paper’s appendix.
Appendix
| Token scalars | Readout parameters | |
|---|---|---|
| 5 | 1,500 | 12 |
| 7 | 2,100 | 16 |
| 256 | 76,800 | 514 |