Multi-Scale Semantic Mapping in Urban Environments via Observation Calibration and Policy Dependence Regularization
Organizations: Harbin Institute of Technology, Shenzhen
Abstract
Semantic mapping is fundamental to embodied navigation, yet existing methods are developed for indoor environments, where objects exhibit relatively limited scale variation and are observed from a restricted range of viewpoints. Urban environments pose substantially greater challenges: agents must map objects ranging from pedestrians to buildings while navigating large spaces with highly diverse viewing distances. These conditions introduce two key difficulties that existing datasets and methods fail to cover. First, object scale and observation distance can be severely mismatched. For example, small objects may be viewed from far away, whereas large objects may be observed at extremely close range, resulting in unreliable observation likelihoods. Second, objects with substantially different sizes and geometries require distinct mapping behaviors, which are difficult to capture with a single shared value estimator. To investigate these challenges, we introduce a large-scale urban semantic mapping dataset featuring realistic city layouts, high-fidelity rendering, and instance-level annotations spanning multiple object scales. We then propose a category-aware likelihood calibration policy that identifies and alleviates unreliable observations according to object category and viewing distance. Because the calibration and motion policies are optimized toward the same mapping objective, they may learn redundant shortcuts and become excessively coupled. We therefore introduce a mutual-information (MI) regularizer that penalizes their estimated representation dependence and encourages complementary behaviors. To better model heterogeneous mapping strategies across object scales, we further employ category-wise value estimators. We formulate their joint optimization as a Pareto optimization problem to mitigate conflicting gradients across categories.
Figures & tables
| Dataset | Type | Platform | Scenes | Scale | Object volume ( ) | Avg. objects/scene | Obj.-level ann. |
| OpenFly [ 19 ] | VLN | UE4 | 21 | - | - | ||
| EmbodiedCity [ 18 ] | VLN | UE5 | 1 | City | - | ||
| UrbanScene 3D [ 30 ] | Map | UE4 | 16 | City | 865.0 | ||
| GLEAM [ 12 ] | Map | Habitat | 1152 | House | - | ||
| MP3D [ 5 ] | Sem-Map | Habitat | 90 | House | 564.6 | ||
| EmbodiedScan [ 41 ] | Sem-Map | Habitat | 5185 | House | 30.9 |
| Method | CCR (%) | |||||
|---|---|---|---|---|---|---|
| CLIP | DINOv3 | CLIP | DINOv3 | CLIP | DINOv3 | |
| Uncertainty | 61.7 0.3 | 62.9 1.9 | 74.7 3.4 | 78.5 2.1 | 95.5 2.1 | 96.2 1.3 |
| Zhang et al. | 52.5 1.5 | 54.7 4.8 | 75.3 4.2 | 79.7 3.3 | 96.5 0.7 | 97.9 0.7 |
| RayFronts | 28.5 5.1 | 27.8 4.4 | 59.8 8.6 | 62.6 6.9 | 91.9 2.2 | 92.2 3.5 |
| ActiveSGM | 57.6 4.8 | 60.2 6.6 | 93.8 2.4 | 95.6 2.2 | 95.7 2.4 | 97.4 1.5 |
| Components | CCR (%) | OCR (%) | Var | |||||
|---|---|---|---|---|---|---|---|---|
| LC | MV | PO | MI | |||||
| – | – | – | – | 82.0 | 97.4 | 99.1 | 92.8 | 59.5 |
| – | – | – | 84.3 | 96.8 | 99.2 | 93.4 | 42.9 | |
| – | – | 50.2 | 86.5 | 99.4 | 78.7 | 433.9 | ||
| – | 92.1 | 97.5 | 99.8 | 96.5 | 10.4 | |||
| – | – | 88.3 | 96.2 | 99.7 | 94.7 | 22.7 | ||
| CCR (%) | OCR (%) | Var | |||
|---|---|---|---|---|---|
| 0.01 | 91.0 | 98.5 | 99.8 | 96.4 | 14.9 |
| 0.05 | 92.7 | 98.5 | 99.8 | 97.0 | 9.6 |
| 0.10 | 93.5 | 99.6 | 99.7 | 97.6 | 8.4 |
| 0.15 | 90.4 | 98.5 | 99.8 | 96.2 | 17.3 |
| 0.20 | 87.8 | 96.1 | 99.9 | 94.6 | 25.4 |
| Method | CCR (%) | OCR (%) | Var | ||
|---|---|---|---|---|---|
| GLEAM-0.2 | 78.9 | 92.4 | 98.4 | 89.9 | 66.2 |
| GLEAM-0.6 | 80.6 | 93.2 | 99.2 | 91.0 | 60.2 |
| GLEAM-1.0 | 83.7 | 93.3 | 98.8 | 91.9 | 38.8 |
| GLEAM-1.4 | 80.3 | 94.5 | 99.4 | 91.4 | 65.5 |
| GLEAM-1.8 | 76.3 | 95.8 | 98.5 | 90.2 | 97.4 |
Appendix figures & tables12 assets
Supplementary material from the paper’s appendix.
Appendix
| You are an intelligent agent to plan a city layput using CityEngine. You will receive a list of scene types and block types, and also a list of available object asset names. You need to design a CGA library for CityEngine. Each CGA should contain these informations: 1. Scene type. 2. Block type. 3. Available object asset names. 4. Distribution functions for each object. 5. Parameters for each distribution function. Do not use scenes, blocks and assets out of the provided lists. Think step by step of the distribution design and output your thought. Output a json file that contains all the information of the CGA library. Scene type: { scene_types } Block type: { block_types } Object assets: { object_assets } |
| You are an intelligent agent to plan a city layput using CityEngine. You will receive a json file of current scene that contains its type with block information, and a list of CGA rules. You need to assign the CGA rules to each block, and assign the parameters of the distribution functions in the rules. Think step by step the type, consider its block type, and the surrounding blocks. Think about how real world objects ditribute and make sure that the parameters are aligned with real world. Scene info path: { scene_info_path } CGA library path: { cga_library_path } |
| Method | mAUC (%) | mIoU (%) | F-1 (%) | |||
|---|---|---|---|---|---|---|
| CLIP | DINOv3 | CLIP | DINOv3 | CLIP | DINOv3 | |
| Uncertainty | 94.3 1.1 | 96.8 0.2 | 70.2 1.8 | 72.6 3.2 | 80.1 4.3 | 82.2 2.8 |
| Zhang et al. | 94.9 2.0 | 96.4 0.9 | 68.0 1.4 | 70.1 0.5 | 78.9 1.8 | 79.8 0.9 |
| RayFronts | 87.8 2.4 | 88.9 2.7 | 55.3 3.4 | 59.1 3.2 | 67.3 2.8 | 69.0 3.5 |
| ActiveSGM | 94.5 0.6 | 96.8 0.2 | 70.6 2.0 | 72.2 2.0 | 79.5 3.1 | 81.4 1.7 |
| GLEAM | 95.8 1.1 | 98.6 0.1 | 88.6 2.1 | 90.9 1.0 | 93.7 0.7 | 95.1 0.5 |
| Components | mAUC (%) | mIoU (%) | F-1 (%) | |||
| LC | MV | PO | MI | |||
| – | – | – | – | 98.5 | 89.8 | 94.4 |
| – | – | – | 99.2 | 92.4 | 95.9 | |
| – | – | 98.9 | 77.4 | 92.8 | ||
| – | 99.4 | 95.1 | 97.5 | |||
| – | – | 99.3 | 94.3 | 97.0 | ||
| mAUC (%) | mIoU (%) | F-1 (%) | |
|---|---|---|---|
| 0.01 | 99.1 | 93.4 | 96.5 |
| 0.05 | 99.3 | 93.4 | 96.5 |
| 0.10 | 99.1 | 95.7 | 97.8 |
| 0.15 | 99.3 | 94.6 | 97.2 |
| 0.20 | 99.0 | 92.4 | 95.9 |
| Method | mAUC (%) | mIoU (%) | F-1 (%) |
|---|---|---|---|
| GLEAM-0.2 | 98.8 | 90.0 | 94.5 |
| GLEAM-0.6 | 98.9 | 90.1 | 94.6 |
| GLEAM-1.0 | 98.5 | 90.9 | 95.1 |
| GLEAM-1.4 | 99.0 | 90.5 | 94.8 |
| GLEAM-1.8 | 98.9 | 88.4 | 93.5 |
| Ours | 99.3 | 94.3 | 97.0 |
| Method | Inference time (ms) |
|---|---|
| Uncertainty | 366.30 |
| Zhang et al. | 401.61 |
| RayFronts | 934.58 |
| ActiveSGM | 578.03 |
| GLEAM | 31.56 |
| Ours | 27.39 |
| Method | Inference time (ms) |
|---|---|
| Uncertainty | 366.30 |
| Zhang et al. | 401.61 |
| RayFronts | 934.58 |
| ActiveSGM | 578.03 |
| GLEAM | 31.56 |
| Ours | 27.39 |
| Method | CCR (%) | OCR (%) | Var | mAUC (%) | mIoU (%) | F-1 (%) | |
|---|---|---|---|---|---|---|---|
| car | building | ||||||
| Uncertainty | 55.9 | 59.5 | 57.7 | 3.3 | 72.0 | 37.8 | 52.8 |
| Zhang et al. | 50.3 | 79.4 | 64.8 | 212.8 | 71.8 | 34.7 | 49.8 |
| RayFronts | 29.5 | 68.5 | 49.0 | 380.6 | 79.6 | 31.3 | 41.2 |
| ActiveSGM | 46.1 | 69.5 | 57.8 | 136.7 | 70.6 | 32.3 | 47.4 |
| GLEAM | 66.8 | 87.5 | 77.1 | 107.8 | 87.2 | 44.6 | 52.4 |
| Configuration | CCR (%) | OCR (%) | Var | mAUC (%) | mIoU (%) | F-1 (%) | ||
|---|---|---|---|---|---|---|---|---|
| LC/LC | 84.3 | 96.8 | 99.2 | 93.4 | 42.9 | 99.2 | 92.4 | 95.9 |
| LC/LC+MI | 86.5 | 96.3 | 99.8 | 94.2 | 31.7 | 99.1 | 92.4 | 95.9 |
| LC+MI/LC | 87.3 | 96.5 | 99.6 | 94.5 | 27.6 | 99.0 | 93.3 | 96.4 |
| LC+MI/LC+MI | 88.3 | 96.2 | 99.7 | 94.7 | 22.7 | 99.3 | 94.3 | 97.0 |
| Metric | LC | LC+MI | Rel. | |
|---|---|---|---|---|
| Max CCA | 0.9910 | 0.9938 | ||
| Mean CCA | 0.0625 | 0.0451 | ||
| Top-10 CCA | 0.6372 | 0.4922 |