cs.ROOct 1, 2026

ColoACT: Multi-Cue Action Chunking for Smooth Autonomous Colon Navigation on a Self-Propelled Endoscopic Robot

Authors: Jian Hu, Shujing He, Leixin Chang, Zongze Li, Ding Huang, Chaoyang Shi, Chengzhi Hu

Organizations: Department of Mechanical and Energy Engineering, Southern University of Science and Technology, Shenzhen, 518055, China · School of Mechanical Engineering and ZJU-UIUC Institute, Zhejiang University, Zhejiang, China · School of Mechanical Engineering, Tianjin University, Tianjin, 300072, China

Abstract

Autonomous colonoscopic navigation can reduce operator burden and the risk of loop formation or tissue trauma, but remains challenging due to deformable anatomy, weak-texture and specular endoscopic visuals, and contact-rich viscoelastic interactions. Existing methods either rely on geometry-driven pipelines, which are efficient and interpretable yet brittle due to manually engineered features and switching logic, or adopt learning-based policies, whose inferred depth/geometry can become temporally inconsistent or overly smooth under weak texture and specular highlights while simulation-trained variants (e.g., deep reinforcement learning) may further suffer from a sim-to-real gap. We propose ColoACT, an autonomous navigation system that integrates an RGB-D-E based Action Chunking Transformer policy (ColoACT policy) for a compact self-propelled Bevel-Gear-Based Endoscopic Robot (BGER). The ColoACT policy augments RGB with estimated relative depth and a gradient-based pseudo-elevation map to enhance fold-ridge saliency and other high-frequency geometric cues, and enables smooth continuous control of the BGER by predicting overlapping action chunks and fusing them via temporal ensembling. In different \textit{ex-vivo} porcine colons (approximately 60 cm), our system achieves success rates of 85.4% and 72.5% in straight and curved segments, respectively, and achieves 70% success in 90-degree turns and 60% in double-bend sequences, with feasibility further demonstrated in challenging triple-bend segments. The project page is available at: https://Adamhu1.github.io/ColoACT/.

Figures & tables

Explore similar work

CardsList
  1. Anatomical Landmark-Guided Deep Reinforcement Learning for Autonomous Gastric Navigation

    May 8, 2026Haoxuan Wu, Sishen Yuan, Haitao Gao +3Robotic Soft EsophagusColonoscopy

  2. BiliVLA: Scene-Aware Vision-Language-Action Model with Reinforcement Learning for Autonomous Biliary Endoscopic Navigation

    Jun 22, 2026Jinsong Lin, Chi Kit Ng, Zhiyong Xiong +8Robotic Soft EsophagusColonoscopy