Organizations: Department of Mechanical and Energy Engineering, Southern University of Science and Technology, Shenzhen, 518055, China · School of Mechanical Engineering and ZJU-UIUC Institute, Zhejiang University, Zhejiang, China · School of Mechanical Engineering, Tianjin University, Tianjin, 300072, China
Autonomous colonoscopic navigation can reduce operator burden and the risk of loop formation or tissue trauma, but remains challenging due to deformable anatomy, weak-texture and specular endoscopic visuals, and contact-rich viscoelastic interactions. Existing methods either rely on geometry-driven pipelines, which are efficient and interpretable yet brittle due to manually engineered features and switching logic, or adopt learning-based policies, whose inferred depth/geometry can become temporally inconsistent or overly smooth under weak texture and specular highlights while simulation-trained variants (e.g., deep reinforcement learning) may further suffer from a sim-to-real gap. We propose ColoACT, an autonomous navigation system that integrates an RGB-D-E based Action Chunking Transformer policy (ColoACT policy) for a compact self-propelled Bevel-Gear-Based Endoscopic Robot (BGER). The ColoACT policy augments RGB with estimated relative depth and a gradient-based pseudo-elevation map to enhance fold-ridge saliency and other high-frequency geometric cues, and enables smooth continuous control of the BGER by predicting overlapping action chunks and fusing them via temporal ensembling. In different \textit{ex-vivo} porcine colons (approximately 60 cm), our system achieves success rates of 85.4% and 72.5% in straight and curved segments, respectively, and achieves 70% success in 90-degree turns and 60% in double-bend sequences, with feasibility further demonstrated in challenging triple-bend segments. The project page is available at: https://Adamhu1.github.io/ColoACT/.
Figures & tables
Fig. 1: System Overview of the proposed ColoACT for Autonomous Navigation of the Bevel-Gear-Based Endoscopic Robot. (a) Framework Overview: The framework integrates a multi-cue perception module (RGB, relative depth, and pseudo-elevation, RGB-D-E) with an ACT-based control policy. Left: ot=[Itrgb,Dt,Et] is formed using Depth Anything V2-Base for relative depth and Scharr gradients for pseudo-elevation. Right: The policy predicts an action chunk at:t+k−1 , which is synthesized via Temporal Ensembling to generate smooth actuation commands at ( vt , ωt ). ( vt , ωt ) represents the normalized linear velocity and angular velocity. (b) The actual navigation of the BGER in the porcine colon. (c) The assembled BGER, with dimensions of 24 × 20 × 34 mm. (d) The overall design of the BGER. (e) Transmission System Design of the BGER.
Fig. 2: Visualization and Stability Analysis. (a) Multi-Cue inputs: RGB, relative depth, and pseudo-elevation map. (b) Temporal standard deviation heatmaps ( N=50 ); the red line marks the sampling row. (c) Dynamic consistency sequence showing stable feature evolution. Cross-sectional profiles at the sampling row of (d) Depth and (e) Elevation. Black lines denote the mean signal, while yellow clusters show raw noise distribution, confirming spatial peak localization.
BC Policy
ACT policy
Input mode
RGB-D-E
RGB
RGB-D
RGB-D-E
Straight
45.8
35.4
58.3
85.4
Curved
57.5
25.0
50.0
72.5
TABLE I: Success rate of the ex-vivo experiment (%)
RGB-D-E based BC
RGB-D-E based ACT
Straight
63.89 ± 7.8877
4.74 ± 0.9312
Curve
58.91 ± 6.5185
4.47 ± 0.4298
TABLE II: The normalized angular Jerk of RGB-D-E based BC policy and ColoACT policy in both straight and curved porcine colon ( a.u. )
Fig. 3: Validation of autonomous navigation in ex-vivo porcine colon. Experiment in (a) Straight colon and (b) Curved colon. The movement of the BGER under weak texture conditions using (c) RGB-D-E based BC and (d) ColoACT policy. Inset photo shows the inside texture of the colon, Scale bar: 5 mm. The normalized linear velocity of (e) RGB-D-E based BC policy and (f) ColoACT policy when facing the weak texture.
Fig. 4: Experiment of ColoACT in complex colon, (a) 90° turns / double-bend and (b) triple-bend, Scale bar: 5 mm. The normalized angular velocity of the following (c) 90° turns / Double-bend and (d) Triple-bend colons.
RGB
RGB-D
RGB-D-E
ADE (a.u.)
0.1187
0.1096
0.0835
FDE (a.u.)
0.2153
0.2076
0.1689
TABLE III: Open-loop trajectory prediction accuracy (ADE/FDE, a.u. ) of the ACT policy across different input configurations; values are means over N = 43 validation samples.
Fig. 5: Temporal ensembling ablation and long-horizon open-loop rollout validation. (a) Temporal ensembling ( k=16 , blue) suppresses jitter compared to the baseline ( k=1 , purple). (b-c) Open-loop trajectory rollouts comparing ACT predictions (dark blue) with demonstrations (Ground truth, light blue) in normalized kinematic space.
Fig. 6: Semantic and attention analysis of policy behavior. (a) t-SNE visualization of the latent space, colored by normalized angular velocity, showing a structured and continuous manifold aligned with steering intent. (b-d) Example observations: RGB image, depth map, and pseudo-elevation cue. (e) Grad-CAM visualization of the visual backbone highlights attention concentrated on the navigable lumen region and relevant anatomical boundaries during turning.
Wireless capsule endoscopy (WCE) enables painless visualization of the gastrointestinal tract, but its diagnostic potential is limited by incomplete mucosal coverage and poor transferability of existing navigation methods across patient anatomies. We propose a transferable, anatomical landmarkguided deep reinforcement learning (AL-DRL) framework for autonomous gastric navigation. Leveraging a lightweight edgecontour-depth fusion module, our policy operates on stable, lowdimensional landmark coordinates rather than high-dimensional video streams, effectively bridging the sim-to-real gap. In simulations across eight patient-derived models, the method achieves over 97% coverage within 50 seconds, significantly outperforming vanilla PPO, SAC, and DQN agents. A two-stage sim-to-real pipeline with an adaptive dynamic programming controller actively mitigates physical disturbances. Ex-vivo experiments demonstrate a mean coverage of 87% and a 53% reduction in procedure time compared with expert manual control.
Haoxuan Wu, Sishen Yuan, Haitao Gao +3
Department of Electronic Engineering, The Chinese University of Hong Kong, Hong Kong SAR, China
Endoscopic retrograde cholangiopancreatography (ERCP) demands precise endoscopic navigation and stable biliary cannulation within a narrow monocular field characterized by specular reflections, partial occlusions, and frequent tissue contact. Although recent robotic systems and vision-based assistance techniques improve operator ergonomics and provide perceptual cues, their performance degrades under pronounced anatomical variability and safety-critical visual artifacts, which hinders reliable autonomy in cannulation-grade procedures. Here, we present BiliVLA, a scene-aware Vision-Language-Action (VLA) framework that formulates biliary endoscopic navigation as an instruction-conditioned visuomotor learning problem. Given an endoscopic observation and a stage-specific language instruction, BiliVLA jointly predicts the target category, a grounded bounding box, and a discrete three-degree-of-freedom (3-DoF) motor command for a continuum endoscope. The proposed framework incorporates scene-aware supervision to improve semantic target consistency and safety-aware recovery supervision to induce conservative retreat behaviors under luminal wall contact. A key component of BiliVLA is a two-stage training paradigm that combines grounding-enhanced supervised fine-tuning (SFT) with Group Relative Policy Optimization (GRPO), thereby improving action reliability and decision consistency during closed-loop navigation. Across three ERCP subtasks, BiliVLA achieves the best overall performance in physical phantom experiments, with a total mIoU of 0.9625, an overall action precision of 91.96%, and an overall success rate (SR) of 84.85%. These results indicate that integrating semantic grounding, scene-aware learning, and reward-guided optimization strengthens perception--action alignment and enables more robust autonomous biliary endoscopic navigation.
Jinsong Lin, Chi Kit Ng, Zhiyong Xiong +8
The Chinese University of Hong Kong, Hong Kong SAR, China. · The Third Affiliated Hospital of Sun Yat-sen University, Guangzhou, China. · University of Cambridge, Cambridge, United Kingdom. +2
Autonomous endoscopic navigation can reduce clinicians' operational burden, yet robust control remains challenging due to tissue deformation, transient occlusions, and rapidly changing viewpoints. Existing learning-based policies typically predict actions from current observations without explicitly modeling future dynamics, limiting their robustness and reliability in safety-critical settings. World Action Models (WAMs) offer a promising alternative by coupling predictive visual dynamics with action generation, but extending them to robotic endoscopy remains challenging due to limited training data, restricted viewpoint diversity, deformable anatomy, and high inference latency. We present EndoWAM, which is, to our knowledge, the first WAM for generalizable robotic endoscopic navigation. EndoWAM introduces future grounding, which predicts task-relevant target regions in future observations from intermediate denoising features of a video world model. Specifically, EndoWAM couples a lightweight diffusion transformer for future target-region prediction with a discrete action expert through a shared predictive representation. This design injects target-aware supervision into predictive dynamics modeling, improving robustness to visual degradation and viewpoint changes while enabling real-time control in a single denoising pass. We further introduce EndoMotion, a robotic endoscopic motion dataset spanning three anatomically distinct procedures: ureteroscopy, esophagoscopy, and endoscopic retrograde cholangiopancreatography (ERCP). EndoWAM consistently outperforms all baselines and alternative grounding strategies, while demonstrating strong zero-shot generalization to unseen viewpoints, environments, and targets. These results establish EndoWAM as a predictive, target-grounded framework for accurate, generalizable, and long-horizon navigation in visually constrained endoscopic environments.