cs.ROJul 17, 2026

Differentiable Reinforcement Learning for Path Tracking by an Agile Fish-Like Robot

Authors: Prashanth ChivkulaKartik LoyaVenkata Ravindhra Reddy VarikutiPhanindra Tallapragada

Abstract

Fish-like swimming has inspired the design of several dozens if not hundreds of bioinspired robots in the last few decades. But the control and motion planning of such robots has been challenging due to the poorly modeled fluid-structure interaction and the nonlinear underactuated dynamics of such robots. While reinforcement learning has allowed significant advances in the context of ground and aerial robots, the lack of a suitable simulation environment with appropriate computational speed and accuracy have prevented similar progress for fish-like robots. We address this two-fold problem by developing a simulation platform that approximates the motion of our fish-like robot with computational efficiency. Then the motion control and path tracking by the robot is performed using PID control where the (variable) gains are learned using back propagation through time and training on a curriculum. The policy learned in the simulation is then applied on the physical platform, demonstrating an excellent match.

Explore similar work

Jun 9, 2026cs.RO

Learning Control as Enabling Layer for Embodied Intelligence Research explored with Soft Robotic Swimming in diverse Flow Speeds

Soft robots are valuable robophysical platforms for studying body-caudal undulatory locomotion, but their compliant bodies are difficult to control precisely under changing hydrodynamic loading. Conventional proportional-integral-derivative (PID) feedback stabilizes periodic undulation in static water, but can accumulate flow-dependent tracking delay and increasing inter-trial variability when environmental flow becomes non-trivial. Here, we evaluate whether augmenting PID control with a Linear Repetitive Learning Estimation Scheme (PID-LRLES) recovers tracking accuracy and repeatability under dynamic flow. The LRLES generalizes classical integral action from constant to periodic, non-constant references, while using a stable transfer-function realization whose poles have negative real parts to avoid the long-term instability issues of classical repetitive control. Closed-loop experiments were carried out in a recirculating flow tank at five bulk flow speeds spanning 0 to 32.6 cm s^-1, using an embedded soft capacitive bending sensor at a 1 kHz control-loop rate. With controller gains tuned once in static water and then held fixed across all conditions, PID-LRLES tracked the periodic bending-envelope reference more closely than the PID baseline and significantly reduced the inter-trial spread of the per-trial RMSE (paired Wilcoxon signed-rank test, p = 1.8 x 10^-4, n = 25). Embedded soft proprioception and cycle-to-cycle learning act as complementary contributors to robustness: the sensor exposes the periodic hydrodynamic bias in body deformation, while the learning term absorbs it over recent oscillation cycles. By reducing flow-dependent control-induced variability, the approach provides an enabling layer for future robophysical studies seeking to isolate the effects of morphology, sensing, and environmental flow on aquatic locomotion.
Fabian Schwab, Federico Allione, Bingcheng Wang +5
Jul 24, 2026cs.RO

Adaptive Undulatory Locomotion of Snake-like Robots in Dynamic Viscous Environments via Deep Reinforcement Learning

This paper demonstrates how deep reinforcement learning (DRL) enables adaptive locomotion of snake-like robots in dynamically changing viscous environments, overcoming the inherent performance limitations of classical predefined control methods. The lack of direct onboard sensors for fluid properties necessitates formulating this task as a partially observable Markov decision process. By employing an asymmetric actor-critic framework, a teacher policy trained using privileged information available only in the physics simulator distills its knowledge into a student policy that relies solely on proprioceptive sensor information. Simulation results across a wide range of dynamic viscosity changes (10710^{-7} to 102m2/s10^{-2} m^2/s) reveal that the DRL agent autonomously acquires non-sinusoidal adaptive gaits. These gaits improve propulsion velocity and transport efficiency, breaking the inherent limits of conventional sinusoidal and kinematic control. The findings establish that implicit environment inference via privileged information distillation is an effective approach to bypass the constraints of classical models under unpredictable fluid dynamics.
Tsuyoshi Kimoto, Akio Yamano, Kohei Honda +1
Jul 29, 2026cs.RO

Reinforcement Learning on Cost-Constrained Quadrupedal Hardware

Deploying learned control policies on low-cost robotic platforms introduces transport latencies and noisy motor feedback that systematically widens the sim-to-real gap. The chasm of simulation to deployment in hardware lies in the delay of the actuator reaching the commanded position. On platforms such as the Mini Pupper 2, a measured >50 ms transport delay transforms the locomotion task from a standard Markov decision process into a partially observable one. In this paper, we take a biologically inspired approach of handling noisy and delayed feedback to close the sim-to-real gap, thereby expanding the capability of reinforcement learning on cost-constrained hardware. Using a low-cost quadrupedal hardware platform, we find that using a forward model of the average actuator delay, paired with a time-aware neural network results in robust locomotion. Additionally, our time-aware neural network learned a central pattern generator (CPG): a self-sustaining rhythmic gait that is robust to +320 ms latency perturbations, mirroring the CPGs found in the spinal cords of vertebrates. We posit that temporal self-organization may be a general strategy for cost-constrained locomotion.
Javier C. Weddington, Bence P. Ölveczky, Stephen A. Baccus