Coupled State-Space Modelling, Control, and Policy Distillation for Hybrid Rigid-Pneumatic Manipulators
Authors: Alan Royce Gabriel Samuel, Pulkit Verma
Organizations: Department of Data Sciences and AI, Indian Institute of Technology Madras, Chennai, India · Department of Computer Science and Engineering, Indian Institute of Technology Madras, Chennai, India
Hybrid manipulators combine motorized rigid joints with pressure-actuated origami segments. Published arms of this kind are controlled with decoupled per-DOF loops, and the cost of this approximation has not been quantified, because the coupled model needed to measure it has not been built. This paper derives such a model for a chain of N alternating revolute joints and Kresling origami segments, including pneumatic chamber dynamics and crease hysteresis. Using the model, we measure the coupling directly and show that its strength varies joint by joint, and that decoupled control loses precisely on the strongly coupled joints while remaining competitive on the one nearly decoupled joint. Coupled model-based controllers track 2.5× tighter than a decoupled PID baseline at lower torque. However, the model predictive controller (MPC) is too slow for real time, and model-free reinforcement learning stalls far below acceptable success rates on a strict settling metric. We therefore distill the MPC into a small neural policy with behavior cloning and DAgger. The distilled policy settles 93-94% of goals with zero collisions, within a few points of its teacher, and runs inside the 5 ms control step where the MPC does not. Where the teacher itself fails, we trace the failure to a limit cycle with the bellows' lightly damped mode, and we remove it by selecting goal postures holdable at low pressure.
Figures & tables
Fig. 1 : Overview: the coupled model, the controllers it supports, and the MPC distilled into a real-time student.
(a)
Holding
Held,
Held,
Peak
Ring after
pressure
default
25 variants
excursion (mm)
1 mm kick (mm)
30 kPa
3/3
2/3–3/3
< 10
1.2–1.6
50 kPa
1/3
1/3 (2/3 in 3)
8–19
3.8–13.2
70 kPa
1/3
0/3–1/3
13–39
7.6–13.4
TABLE I : Hold test at N=3 : MPC regulates at a goal’s static equilibrium for 500 steps. Held means never leaving 10 mm, 3 goals per bin, over 25 controller variants. Last column: ring after a 1 mm kick, outer QP removed.
Controller
θ1 (rad)
θ2 (rad)
θ3 (rad)
RMS τ (N m)
PID (decoupled)
0.073
0.013
0.064
4.83
LQR
0.020
0.029
0.017
1.32
LQR (decoupled gain)
0.023
0.021
0.017
1.32
MPC
0.031
0.025
0.014
1.32
TABLE II : Tracking on the single-target benchmark without the disturbance estimator (RQ3): RMS joint error per joint (yaw, pitch, yaw) and the worst joint’s RMS torque
Controller
Success
Coll.
Steps
Model-free SAC + HER (plateau)
46.8%
1.2%
Behavior cloning only
55.4%
16.0%
136
DAgger round 1 (3 seeds)
90.3%
0.6%
110
DAgger round 2 (2 seeds)
93.7%
0.0%
106
MPC teacher (demonstration goals)
91.4%
0.0%
108
MPC teacher / student, identical goals
95.0 / 92.4%
0.0%
TABLE III : Distillation at N=2 (settled, 10 mm, 500 episodes per seed; steps: mean steps to settle). DAgger round 2 pools two seeds at 93.2 and 94.2.
Reinforcement learning-based control policies have been frequently demonstrated to be more effective than analytical techniques for many manipulation tasks. Commonly, these methods learn neural control policies that predict end-effector pose changes directly from observed state information. For tasks like inserting delicate connectors which induce force constraints, pose-based policies have limited explicit control over force and rely on carefully tuned low-level controllers to avoid executing damaging actions. In this work, we present hybrid position-force control policies that learn to dynamically select when to use force or position control in each control dimension. To improve learning efficiency of these policies, we introduce Mode-Aware Training for Contact Handling (MATCH) which adjusts policy action probabilities to explicitly mirror the mode selection behavior in hybrid control. We validate MATCH's learned policy effectiveness using fragile peg-in-hole tasks under extreme localization uncertainty. We find MATCH substantially outperforms pose-control policies -- solving these tasks with up to 10% higher success rates and 5x fewer peg breaks than pose-only policies under common types of state estimation error. MATCH also demonstrates data efficiency equal to pose-control policies, despite learning in a larger and more complex action space. In over 1600 sim-to-real experiments, we find MATCH succeeds twice as often as pose policies in high noise settings (33% vs.~68%) and applies ~30% less force on average compared to variable impedance policies on a Franka FR3 in laboratory conditions.
Manipulation in confined environments, such as threading a manipulator through narrow apertures, remains a fundamental challenge, especially for conventional rigid robots. Hybrid rigid-soft manipulators offer promise but face two compounding planning challenges: backbone shapes feasible in free space become infeasible under environmental contact, and planning rigid and soft segments independently ignores their kinematic coupling. We present THREAD, the first diffusion-based trajectory planner for hybrid manipulation, learning a generative prior over physically realizable backbone trajectories conditioned on local environment geometry, with physics-inspired losses encoding curvature, smoothness, and collision constraints jointly across both segments. Trained in simulation, THREAD achieves 92.4% task success with 5x fewer collisions than the strongest baseline. We show cross-embodiment real-world transfer with minimal online updates, successfully threading through apertures as small as 1.3x the soft segment diameter.
While Model Predictive Control (MPC) provides strong stability and robustness, it imposes a significant computational burden on real-time systems. This paper investigates the application of Behavior Cloning to approximate MPC policies for the real-time control of a 3-degree-of-freedom robotic manipulator. We present a baseline controller combining Inverse Kinematics with MPC and evaluate neural network architectures, ranging from classical regression algorithms to deep learning models including Deep MLPs and RNNs, to derive computationally efficient surrogate policies. We analyze generalization capabilities, stability considerations, and the trade-offs inherent in different architectural choices. Our empirical study employs both online and offline evaluations to assess performance regarding accuracy, computational efficiency, and fidelity to the original MPC policy. Our results demonstrate that Behavior Cloning can effectively reduce the computational burden of MPC policies for 3-DOF robotic manipulators, achieving a 3x reduction in inference latency with a 84.98% success rate under relaxed tolerances. Notably, we find that static architectures outperform temporal variants, confirming the sufficiency of instantaneous state observations for this task. However, we observe a precision gap under strict tolerances, which suggest that while Behavior Cloning captures the global optimal trajectory, further research is needed to minimize terminal steady-state error.