ExcavaTwin: Training-Free Geometry-Guided Semantic Elevation Mapping for Autonomous Excavation
Organizations: The Smart Manufacturing Thrust, The Hong Kong University of Science and Technology (Guangzhou) · Department of Electronic and Electrical Engineering, Southern University of Science and Technology. · Capstone Technology Co., Ltd., Shenzhen, China.
Abstract
Autonomous excavation requires a spatial representation that jointly captures terrain geometry and task-relevant semantics. Existing excavation mapping is largely elevation-centric, while generic semantic models remain unstable in unstructured outdoor scenes. We present ExcavaTwin, a pure-vision geometry-guided semantic elevation mapping framework without excavation-specific training. Given multi-view RGB images, the framework: 1) reconstructs scene geometry and semantic observations using frozen vision models; 2) derives terrain and non-terrain geometric support; 3) performs geometry-constrained multi-view semantic fusion to suppress implausible predictions and recover incomplete observations; and 4) projects the fused state into a task-oriented semantic elevation map. Experiments on public datasets and real excavation scenes demonstrate reliable geometric and semantic perception. In real excavation, the system achieved an average update interval of approximately 1.4 s and a mean elevation error of 12.74cm in dynamically modified regions. Larger errors mainly occur during rapid terrain changes and transient visual disturbances caused by machine motion.
Figures & tables
| Dataset | Source | Key scale | Visual input | Geometric Reference | Evaluation role |
|---|---|---|---|---|---|
| RELLIS-3D [ 8 ] | Public | 13,556 LiDAR scans; 6,235 RGB images; 20 semantic classes | RGB video | LiDAR-based terrain reference | Terrain elevation |
| Objectron | Public | 14,819 annotated video sequences; 9 categories; 180 videos sampled | RGB video | Object-level metric reference | Object dimension recovery |
| GOOSE-Ex [ 19 ] | Public | Selected outdoor sequences with semantic annotations | RGB video | – | Semantic mapping |
| OCES | Self-collected | 25 video sequences; 5 object categories; 15k RGB frames; 3,500 LiDAR scans | RGB video | LiDAR-based terrain reference + object dimensions | Object dimension recovery; terrain elevation; semantic mapping |
| Task | Dataset | MAE (cm) | RMSE (cm) | Rel. Error (%) |
|---|---|---|---|---|
| Object Dimension | Objectron | 2.63 | 4.55 | 9.01 |
| OCES | 6.23 | 8.94 | 10.43 | |
| Terrain Elevation | RELLIS-3D | 28.30 | 32.14 | – |
| OCES | 12.61 | 15.32 | – |
| GOOSE-Ex | OCES | ||||
|---|---|---|---|---|---|
| Backbone | Method | mIoU (%) | Coverage (%) | mIoU (%) | Coverage (%) |
| YOLOE | Backbone | 3.34 | 2.57 | 6.53 | 2.70 |
| YOLOE | + Ours | 20.77 | 90.52 | 34.57 | 80.40 |
| Mask2Former | Backbone | 21.40 | 94.26 | 34.71 | 91.85 |
| Mask2Former | + Ours | 38.03 | 96.15 | 53.19 | 99.28 |
| GOOSE-Ex | OCES | |||
| Configuration | mIoU | Cov. | mIoU | Cov. |
| YOLOE backbone | 3.34 | 2.57 | 6.53 | 2.70 |
| w/o Local geometry | 9.12 | 70.46 | 21.82 | 64.13 |
| w/o Terrain continuity | 17.84 | 81.28 | 29.63 | 79.72 |
| w/o Object support | 18.40 | 89.07 | 30.28 | 81.94 |
| w/o Spatial consistency | 10.36 | 63.18 | 17.92 | 50.47 |