Repeated excavation continuously reshapes pile geometry, requiring an autonomous excavator to adapt its digging targets and coordinate motion across successive excavation cycles. We present a learning-based framework for continuous autonomous excavation that integrates terrain-aware target selection with reinforcement- and imitation-learning controllers. The framework separates target-conditioned motion from local digging: a shared task-conditioned RL policy controls waypoint-guided approach and loaded transport, while an IL policy learns vision-based digging and lifting from expert demonstrations. Digging targets are selected from LiDAR elevation maps and converted into bucket-tip waypoints for motion control. The control architecture coordinates the learned policies and deterministic unloading through a shared motion interface. The complete system is deployed on a scaled hydraulic excavator with multimodal sensing and closed-loop actuator control. Offline replay and physical experiments demonstrate more consistent target selection, shorter local motion time, and increased payload compared with the respective baselines. The learned digging policy achieves a mean payload of 6.52 kg per completed cycle, compared with 2.68 kg for Fixed Dig. Three five-scoop runs further demonstrate consecutive autonomous excavation under continuously changing pile geometry.
Figures & tables
Fig. 1: Hardware Platform and System Integration. We robotize a teleoperated scaled hydraulic excavator by integrating multimodal sensing, actuator feedback, Jetson Orin NX policy inference, and STM32-based low-level control. Dimensions show overall length and height in the photographed pose.
Fig. 2: Control Architecture. The Temporal Target Selector supplies targets for waypoint planning and shared RL control of approach and loaded transport. ACT uses RGB and proprioception for digging and lifting. Both policies and Fixed Dump share a motion interface. Solid arrows show data flow; the dashed arrow closes the cycle through renewed terrain observation.
Fig. 3: Repeated excavation workflow. The RL Tracker handles approach and loaded transport; ACT performs digging and lifting; Fixed Dump unloads. The Temporal Target Selector selects the next target, followed by waypoint planning before the next approach. The inset in panel 6 shows a captured visualization of the Temporal Target Selector with the selected target marked by a magenta sphere.
Fig. 4: Offline target selections over middle-frame elevation maps. Rows show three soil resets; columns show initial and manually depressed surfaces. Highest, Spatial, and Temporal denote highest-point selection, Spatial Only, and the Temporal Target Selector; markers show all valid targets. Shared colors encode machine-frame surface Z ; blank cells lack valid height.
Sequence
Frames
H jumps
S jumps
F jumps
F outputs
R1 I
161
61
50
0
159
R1 C
155
44
98
0
151
R2 I
176
67
57
0
174
R2 C
204
48
82
0
202
R3 I
186
46
133
0
184
R3 C
180
34
103
1
178
TABLE I: Static-terrain target selection. H: global highest point; S: Spatial Only; F: Temporal Target Selector. Jumps are adjacent-valid-frame XY changes above 0.1 m. I/C denote initial/depressed surfaces within a reset.
Metric
RL Tracker
DLS
Position success
5/5
5/5
Tracking time (s)
1.04±0.54
1.93±0.11
Terminal position error (cm)
1.62
1.77
Absolute pitch change ( ∘ )
2.53
3.56
TABLE II: Empty-bucket 10-cm motion at 45∘ upward, five trials per controller. Time is mean ± sample SD; other continuous metrics are means. Position errors use the machine’s kinematic estimate.
Pair
Surface
Temporal Target Selector
Highest
Difference
P1
F
6.05
6.05
0.00
P2
F
6.85
6.75
+0.10
P3
F
6.55
5.50
+1.05
P4
D
6.70
6.25
+0.45
P5
D
6.90
6.90
0.00
P6
D
7.05
6.70
+0.35
TABLE III: Target-selection comparison: net mass (kg), one cycle per method in each pair. F: leveled soil; D: central depression.
Run
Scoops
Net mass (kg)
Duration (s)
S1
5
31.75
204.40
S2
5
29.90
235.14
S3
5
33.10
193.95
Total
15
94.75
—
TABLE IV: Five-scoop runs, weighed per run. Duration includes startup and finalization.
Autonomous excavator control is challenged by coupled kinematics, actuation lag, and uncertainty. We propose imitation learning and adaptive Cartesian tracking (IL-ACT), a novel motion control framework for a 30-ton-class excavator. An anchored, 14-input imitation policy pretrained on operator demonstrations generates nominal joint rates; adaptive Cartesian feedback and gated gain/bias estimation correct these commands before a stopping-distance governor constrains joint-reference generation. Simscape evaluation covers 100 sequential goals and spiral, figure-eight, and rounded-raster tracking, including 88 additional runs across three training seeds, two initializations, and speeds, under hydraulic response and sensing conditions. Compared with Teacher+ACT, IL-ACT completes all goals with shorter duration and lower terminal errors under both response conditions. Telemetry-initialized IL-ACT lowers RMSE in all 24 figure-eight and rounded-raster seed comparisons and lowers additional-load spiral mean RMSE by approximately 29%. Original spiral RMSE also improves over IL-only and PID. Under a shared sensor-noise realization, telemetry-initialized IL-ACT achieves 27.67% lower mean RMSE than Teacher+ACT; enabling estimation reduces mean RMSE by 22.44% relative to the frozen estimator. Pretrained-weight effects remain mixed, and the original teacher comparison exhibits a spiral RMSE--maximum-error tradeoff. Analysis establishes bounded adaptive states and Cartesian feedback, with reference admissibility conditional on governor feasibility.
Mehdi Heydari Shahna, Seihun Kim, Soyi Jung +3
Faculty of Engineering and Natural Sciences, Tampere University, Tampere, Finland · Department of Defense Convergence Technology, Korea University, Seoul, Republic of Korea · Department of Electrical and Computer Engineering, Korea University, Seoul, Republic of Korea +2
Earthmoving tasks such as excavation, backfilling, or embankment construction require deliberate repositioning of deformable soil. For these tasks, human operators use all shovel faces, while autonomous systems so far are limited to excavation and dumping. Current methods often rely on heuristic models but do not incorporate soil mechanics. We address this shortcoming by using Reinforcement Learning in a GPU-parallelized Material Point Method particle simulation. Our controllers are conditioned on material state such as shape and compactness, enabling skills that use multiple contact faces of the tool and displace material both inside and outside of the shovel. To use the same learned weights across machines, our policies operate in a normalized end-effector space and are deployed through a calibrated machine interface. We evaluate this calibrated transfer on an 11.5t hydraulic excavator and a 500g tabletop robot. We validate performance through autonomous construction of a 42m long, 2.1m high embankment in 45min, executing 201 individual policy strokes without failure, retry, or operator intervention. In a direct comparison, the autonomous controller matches an expert operator's progression speed and produces a higher, more consistent embankment. Additional qualitative backfilling and compaction experiments demonstrate the material-state awareness and calibrated transfer across machines.
Lennart Werner, Pol Eyschen, Sean Costello +3
Robotic Systems Lab, ETH Zürich, Leonhardstrasse 21, 8092 Zürich, Switzerland · Hexagon Innovation Hub GmbH, Heinrich-Wild-Strasse 201, 9435 Heerbrugg, Switzerland
Autonomous obstacle removal from the ground is an important earthwork task, but this is difficult to automate because an excavator must adapt its excavation trajectories over repeated cycles as soil-obstacle conditions change. Learning such state-dependent behavior requires a training environment that reproduces accumulated soil-obstacle interactions, including contact states, terrain deformation, and obstacle visibility. Accordingly, particle-based simulation is suitable for the relevant policy learning. However, particle simulation is computationally expensive, and repeated excavation cycles further increase the learning cost. We observe that the burial condition of an obstacle governs both task difficulty and simulation cost: deeper burial makes obstacle removal harder while also requiring more particles for accurate simulation. This observation motivates a burial-conditioned curriculum learning strategy. We propose a time-efficient sim-to-real policy learning framework in which the policy observes terrain and obstacle information from RGB-D measurements and then outputs a parameterized excavation trajectory; in this process, the simulator reproduces in a real-world excavator the same observation-action interface it uses under controllable burial conditions. The curriculum begins with shallow burial conditions and progressively increases burial depth while adjusting particle count, thus simultaneously controlling task difficulty and simulation cost. Experiments show that the proposed framework successfully learns an effective obstacle-removal policy, whereas baseline methods fail even after a full week of training. The proposed curriculum achieves effective performance within three days and achieves successful transfer to a real 12-ton excavator operating on open ground with various steel obstacles, thus demonstrating robust obstacle removal.
Yuki Kadokawa, Sandro M. Alcantara Tacora, Taro Abe +4
Nara Institute of Science and Technology · Public Works Research Institute, Ibaraki 300-2621, Japan