Coupled State-Space Modelling, Control, and Policy Distillation for Hybrid Rigid-Pneumatic Manipulators
Authors: Alan Royce Gabriel Samuel, Pulkit Verma
Organizations: Department of Data Sciences and AI, Indian Institute of Technology Madras, Chennai, India · Department of Computer Science and Engineering, Indian Institute of Technology Madras, Chennai, India
Hybrid manipulators combine motorized rigid joints with pressure-actuated origami segments. Published arms of this kind are controlled with decoupled per-DOF loops, and the cost of this approximation has not been quantified, because the coupled model needed to measure it has not been built. This paper derives such a model for a chain of N alternating revolute joints and Kresling origami segments, including pneumatic chamber dynamics and crease hysteresis. Using the model, we measure the coupling directly and show that its strength varies joint by joint, and that decoupled control loses precisely on the strongly coupled joints while remaining competitive on the one nearly decoupled joint. Coupled model-based controllers track 2.5× tighter than a decoupled PID baseline at lower torque. However, the model predictive controller (MPC) is too slow for real time, and model-free reinforcement learning stalls far below acceptable success rates on a strict settling metric. We therefore distill the MPC into a small neural policy with behavior cloning and DAgger. The distilled policy settles 93-94% of goals with zero collisions, within a few points of its teacher, and runs inside the 5 ms control step where the MPC does not. Where the teacher itself fails, we trace the failure to a limit cycle with the bellows' lightly damped mode, and we remove it by selecting goal postures holdable at low pressure.
Figures & tables
Fig. 1 : Overview: the coupled model, the controllers it supports, and the MPC distilled into a real-time student.
(a)
Holding
Held,
Held,
Peak
Ring after
pressure
default
25 variants
excursion (mm)
1 mm kick (mm)
30 kPa
3/3
2/3–3/3
< 10
1.2–1.6
50 kPa
1/3
1/3 (2/3 in 3)
8–19
3.8–13.2
70 kPa
1/3
0/3–1/3
13–39
7.6–13.4
TABLE I : Hold test at N=3 : MPC regulates at a goal’s static equilibrium for 500 steps. Held means never leaving 10 mm, 3 goals per bin, over 25 controller variants. Last column: ring after a 1 mm kick, outer QP removed.
Controller
θ1 (rad)
θ2 (rad)
θ3 (rad)
RMS τ (N m)
PID (decoupled)
0.073
0.013
0.064
4.83
LQR
0.020
0.029
0.017
1.32
LQR (decoupled gain)
0.023
0.021
0.017
1.32
MPC
0.031
0.025
0.014
1.32
TABLE II : Tracking on the single-target benchmark without the disturbance estimator (RQ3): RMS joint error per joint (yaw, pitch, yaw) and the worst joint’s RMS torque
Controller
Success
Coll.
Steps
Model-free SAC + HER (plateau)
46.8%
1.2%
Behavior cloning only
55.4%
16.0%
136
DAgger round 1 (3 seeds)
90.3%
0.6%
110
DAgger round 2 (2 seeds)
93.7%
0.0%
106
MPC teacher (demonstration goals)
91.4%
0.0%
108
MPC teacher / student, identical goals
95.0 / 92.4%
0.0%
TABLE III : Distillation at N=2 (settled, 10 mm, 500 episodes per seed; steps: mean steps to settle). DAgger round 2 pools two seeds at 93.2 and 94.2.