Towards Spatial Perception for Heterogeneous Robot Collaboration in Subterranean Mining Environments
Authors: Mario Alberto Valdes Saucedo, Akash Patel, Christoforos Kanellakis, George Nikolakopoulos
Organizations: Robotics and Artificial Intelligence Group, Department of Computer Science, Electrical and Space Engineering, Luleå University of Technology, 971 87 Luleå, Sweden
The autonomous extraction of deep mineral deposits in abandoned underground mines is fundamentally a multi-agent integration problem. No single platform simultaneously offers the mobility to traverse kilometers of degraded drifts and the sensing payload required to characterize an ore body. This article presents the onboard perception pipeline that bridges two heterogeneous agents within the PERSEPHONE autonomous mining mission. Which consist of a lightweight Explorer robot that maps an unknown mine and generates a 3D scene graph of inspection targets, by running a zero-shot, vision-language semantic segmentation stack that detects mineral deposits directly from natural-language prompts. The map and the graph are then handed to a second Inspector robot, which carries an advanced sensing payload and uses them to plan close-range inspection viewpoints. We detail the complete pipeline, with emphasis on the geometric abstraction that turns raw detections into actionable inspection targets, spanning per-view bounding-box generation, cross-view box merging, plane fitting, and polygon extraction, and we report an extensive field validation in a subterranean test facility and in an active magnesite mine, covering both iron-vein and magnesite mineralization under realistic, perceptually degraded conditions.
Figures & tables
Fig. 1 : Heterogeneous two-robot mining mission. The Explorer maps the mine while running the proposed onboard open-set mineral perception pipeline, it then transmits a PCD map and a set of oriented mineral polygons to the Inspector , which plans close-range inspection along each polygon’s normal.
Fig. 2 : Onboard perception pipeline. The Explorer transforms its raw sensor streams (top: RGB image, language prompt, LiDAR cloud, and odometry) into a compact 3D scene graph of inspection targets through five stages. See – zero-shot segmentation with CLIPSeg detects the prompted mineral in the image; Locate – the labeled mask is fused with the synchronized LiDAR scan to ground each detection in 3D; Track – detections are merged and deduplicated across the trajectory into persistent deposit regions; Model – each region is abstracted as an oriented planar polygon encoding position, extent, and surface normal; Plan – the polygons are organized into a 3D scene graph and transmitted to the Inspector .
Fig. 3 : Iron-vein results (LTU SubT facility). Right/center : the reconstructed mine point cloud overlaid with the onboard-generated 3D scene graph, in which a root Mine node (magenta) connects to the mineral area clusters (orange), each grouping the per-deposit mineral bounding boxes (green). Top-left insets : the oriented inspection polygons (green) and their surface normals (red), fitted to representative deposits, the cloud is colored by per-point detection confidence. Bottom : eight representative onboard frames with the per-image CLIPSeg detection-confidence heat-maps from which the mineral masks are extracted.
Fig. 4 : Magnesite results (Grecian Magnesite, Koutizi). Top-left : the 3D scene graph linking the root Mine node (magenta) to its mineral area clusters (orange) and mineral bounding boxes (green). Top-right : the resulting inspection targets , i.e. the oriented polygons (green) and surface normals (red) over the confidence-colored point cloud. Bottom : per-image CLIPSeg detection-confidence heat-maps.
Fig. 5 : Quantitative analysis. (a) Per-point CLIPSeg detection confidence cleanly separates the detected mineral ( cˉ≈0.60 ) from the host rock ( cˉ≈0.02 ), with the segmentation threshold τ (dashed) lying in the valley between the two modes. (b) Over the run the ∼38 k raw mineral points (orange) are condensed by the cross-view clustering of Sec. III-C into a few dozen persistent deposits and inspection polygons (blue/green), a handover roughly two orders of magnitude smaller than the raw semantic cloud. (c) Distribution of inspection-polygon areas across deposits, spanning compact pockets to larger vein faces. (d) RANSAC plane-fit inlier ratio versus deposit size, color corresponds to polygon area, compact deposits are well approximated by a single plane, while the largest deposits are less planar and are better described by multiple polygons.
College of Artificial Intelligence, Harbin Institute of Technology (Shenzhen), Shenzhen 518055, China · Department of Mechanical and Automation Engineering, The Chinese University of Hong Kong, Hongkong 999077, China
School of Minerals and Energy Resources Engineering, University of New South Wales, Sydney, NSW, Australia · School of Science, Engineering and Digital Technologies, University of Southern Queensland, Toowoomba, QLD, Australia.