cs.ROSep 28, 2026

ExcavaTwin: Training-Free Geometry-Guided Semantic Elevation Mapping for Autonomous Excavation

Authors: Yu Deng, Lingshan Zeng, Tong Hu, Rushi Dai

Organizations: The Smart Manufacturing Thrust, The Hong Kong University of Science and Technology (Guangzhou) · Department of Electronic and Electrical Engineering, Southern University of Science and Technology. · Capstone Technology Co., Ltd., Shenzhen, China.

Abstract

Autonomous excavation requires a spatial representation that jointly captures terrain geometry and task-relevant semantics. Existing excavation mapping is largely elevation-centric, while generic semantic models remain unstable in unstructured outdoor scenes. We present ExcavaTwin, a pure-vision geometry-guided semantic elevation mapping framework without excavation-specific training. Given multi-view RGB images, the framework: 1) reconstructs scene geometry and semantic observations using frozen vision models; 2) derives terrain and non-terrain geometric support; 3) performs geometry-constrained multi-view semantic fusion to suppress implausible predictions and recover incomplete observations; and 4) projects the fused state into a task-oriented semantic elevation map. Experiments on public datasets and real excavation scenes demonstrate reliable geometric and semantic perception. In real excavation, the system achieved an average update interval of approximately 1.4 s and a mean elevation error of 12.74cm in dynamically modified regions. Larger errors mainly occur during rapid terrain changes and transient visual disturbances caused by machine motion.

Figures & tables

Explore similar work

May 22, 2026cs.RO

Semantic-Aware Guided Drone Exploration for Language-Conditioned 3D Indoor Mapping

We present Semantic-Aware Guided Exploration, SAGE, a system for open-vocabulary exploration in unknown 3D indoor environments that preserves coverage-oriented behavior while allowing semantic cues to reprioritize frontier selection. Building on the FALCON volumetric explorer, SAGE integrates Contrastive Language-Image Pre-training (CLIP) via four key components: object-centric embedding storage, a temporal cache that projects recent observations onto the free-unknown boundary, object frontiers for high-similarity detections, and a unified semantic-geometric planning cost. This cost function bounds semantic reweighting influence, ensuring frontiers are prioritized without sacrificing total coverage. In Matterport3D-based simulations, SAGE outperforms FALCON and a semantic-only ablation in object discovery across map-query pairs. Compared to Finding Things in the Unknown (FTU), SAGE completes exploration 9.0 to 25.9 times faster across the nine shared map-query pairs, achieving a mean speedup of 13.7. Furthermore, SAGE achieves substantially higher volumetric throughput than FTU. Finally, we deploy SAGE in five real-world flights in two environments on a Modal AI Starling 2 quadrotor with onboard sensing and planning, and offboard CLIP inference. Comparing SAGE and FALCON, we find that while FALCON results in faster exploration and shorter mapping trajectories, SAGE outperforms FALCON in terms of object discovery.
Jun 24, 2026cs.RO

RoboAtlas: Contextual Active SLAM

We present RoboAtlas, a contextual Active SLAM framework that adaptively balances geometric exploration and semantic reasoning using a scalable 3D semantic mapping system, OpenRoboVox. RoboAtlas integrates frontier exploration, global semantic-map reasoning, and egocentric VLM-based reasoning through a contextual multi-armed bandit that transitions from exploration to semantically guided navigation as scene understanding improves. We evaluate the system in simulation and on a Unitree Go2 robot in large-scale real-world environments exceeding 1800 m2 with approx. 30k mapped semantic instances, achieving a 100% task success rate. On the GOAT-Bench "Val Unseen" benchmark, RoboAtlas achieves state-of-the-art performance with highest reported success rate (SR) of 90.6%, using GPT-4o, improving over the strongest prior baseline by 17.8 percentage points in SR. Using the much smaller Qwen2.5-VL-7B model, it still achieves 88.8% SR, outperforming all baselines using GPT-4o in SR, and revealing the importance of the information gained by our semantic mapping framework over simply replacing the underlying foundation model. The results demonstrate that grounding foundation models with large-scale 3D semantic maps enables robust and efficient contextual Active SLAM.
Oct 7, 2026cs.CV

ActiveLang: Active Open-Vocabulary 3D Mapping with Semantic-Uncertainty-Guided Exploration

As robots increasingly assist humans with diverse tasks, they need both geometric and semantic understanding of their surroundings. Moreover, robots often operate in unfamiliar environments and take on new tasks without knowing the relevant concepts ahead of time. This motivates language-annotated 3D maps that support open-vocabulary scene understanding and human-robot interaction. We introduce ActiveLang, an autonomous system for active open-vocabulary 3D mapping with semantic-uncertainty-guided exploration. ActiveLang performs online language-feature adaptation on a compact dual-Gaussian representation to jointly reconstruct scene geometry, appearance, and open-vocabulary semantics with modest memory overhead. Its planner efficiently selects informative viewpoints, enabling effective mapping with fewer observations and lower computational cost. Experiments on Replica and ScanNet++ demonstrate substantial improvements in 2D and 3D open-vocabulary segmentation over both online and offline baselines, highlighting that actively exploring scenes builds language-annotated 3D maps more efficiently.