Organizations: Department of Mechanical and Energy Engineering, Southern University of Science and Technology, Shenzhen, 518055, China · School of Mechanical Engineering and ZJU-UIUC Institute, Zhejiang University, Zhejiang, China · School of Mechanical Engineering, Tianjin University, Tianjin, 300072, China
Autonomous colonoscopic navigation can reduce operator burden and the risk of loop formation or tissue trauma, but remains challenging due to deformable anatomy, weak-texture and specular endoscopic visuals, and contact-rich viscoelastic interactions. Existing methods either rely on geometry-driven pipelines, which are efficient and interpretable yet brittle due to manually engineered features and switching logic, or adopt learning-based policies, whose inferred depth/geometry can become temporally inconsistent or overly smooth under weak texture and specular highlights while simulation-trained variants (e.g., deep reinforcement learning) may further suffer from a sim-to-real gap. We propose ColoACT, an autonomous navigation system that integrates an RGB-D-E based Action Chunking Transformer policy (ColoACT policy) for a compact self-propelled Bevel-Gear-Based Endoscopic Robot (BGER). The ColoACT policy augments RGB with estimated relative depth and a gradient-based pseudo-elevation map to enhance fold-ridge saliency and other high-frequency geometric cues, and enables smooth continuous control of the BGER by predicting overlapping action chunks and fusing them via temporal ensembling. In different \textit{ex-vivo} porcine colons (approximately 60 cm), our system achieves success rates of 85.4% and 72.5% in straight and curved segments, respectively, and achieves 70% success in 90-degree turns and 60% in double-bend sequences, with feasibility further demonstrated in challenging triple-bend segments. The project page is available at: https://Adamhu1.github.io/ColoACT/.
Figures & tables
Fig. 1: System Overview of the proposed ColoACT for Autonomous Navigation of the Bevel-Gear-Based Endoscopic Robot. (a) Framework Overview: The framework integrates a multi-cue perception module (RGB, relative depth, and pseudo-elevation, RGB-D-E) with an ACT-based control policy. Left: ot=[Itrgb,Dt,Et] is formed using Depth Anything V2-Base for relative depth and Scharr gradients for pseudo-elevation. Right: The policy predicts an action chunk at:t+k−1 , which is synthesized via Temporal Ensembling to generate smooth actuation commands at ( vt , ωt ). ( vt , ωt ) represents the normalized linear velocity and angular velocity. (b) The actual navigation of the BGER in the porcine colon. (c) The assembled BGER, with dimensions of 24 × 20 × 34 mm. (d) The overall design of the BGER. (e) Transmission System Design of the BGER.
Fig. 2: Visualization and Stability Analysis. (a) Multi-Cue inputs: RGB, relative depth, and pseudo-elevation map. (b) Temporal standard deviation heatmaps ( N=50 ); the red line marks the sampling row. (c) Dynamic consistency sequence showing stable feature evolution. Cross-sectional profiles at the sampling row of (d) Depth and (e) Elevation. Black lines denote the mean signal, while yellow clusters show raw noise distribution, confirming spatial peak localization.
BC Policy
ACT policy
Input mode
RGB-D-E
RGB
RGB-D
RGB-D-E
Straight
45.8
35.4
58.3
85.4
Curved
57.5
25.0
50.0
72.5
TABLE I: Success rate of the ex-vivo experiment (%)
RGB-D-E based BC
RGB-D-E based ACT
Straight
63.89 ± 7.8877
4.74 ± 0.9312
Curve
58.91 ± 6.5185
4.47 ± 0.4298
TABLE II: The normalized angular Jerk of RGB-D-E based BC policy and ColoACT policy in both straight and curved porcine colon ( a.u. )
Fig. 3: Validation of autonomous navigation in ex-vivo porcine colon. Experiment in (a) Straight colon and (b) Curved colon. The movement of the BGER under weak texture conditions using (c) RGB-D-E based BC and (d) ColoACT policy. Inset photo shows the inside texture of the colon, Scale bar: 5 mm. The normalized linear velocity of (e) RGB-D-E based BC policy and (f) ColoACT policy when facing the weak texture.
Fig. 4: Experiment of ColoACT in complex colon, (a) 90° turns / double-bend and (b) triple-bend, Scale bar: 5 mm. The normalized angular velocity of the following (c) 90° turns / Double-bend and (d) Triple-bend colons.
RGB
RGB-D
RGB-D-E
ADE (a.u.)
0.1187
0.1096
0.0835
FDE (a.u.)
0.2153
0.2076
0.1689
TABLE III: Open-loop trajectory prediction accuracy (ADE/FDE, a.u. ) of the ACT policy across different input configurations; values are means over N = 43 validation samples.
Fig. 5: Temporal ensembling ablation and long-horizon open-loop rollout validation. (a) Temporal ensembling ( k=16 , blue) suppresses jitter compared to the baseline ( k=1 , purple). (b-c) Open-loop trajectory rollouts comparing ACT predictions (dark blue) with demonstrations (Ground truth, light blue) in normalized kinematic space.
Fig. 6: Semantic and attention analysis of policy behavior. (a) t-SNE visualization of the latent space, colored by normalized angular velocity, showing a structured and continuous manifold aligned with steering intent. (b-d) Example observations: RGB image, depth map, and pseudo-elevation cue. (e) Grad-CAM visualization of the visual backbone highlights attention concentrated on the navigable lumen region and relevant anatomical boundaries during turning.
The Chinese University of Hong Kong, Hong Kong SAR, China. · The Third Affiliated Hospital of Sun Yat-sen University, Guangzhou, China. · University of Cambridge, Cambridge, United Kingdom. +2