Physics Residual Dynamics and Reduced Order Whole-Body Planning for Obstacle Aware Human Robot Cloth CoTransportation
Authors: Moein Forouhar, Kosar Behnia, Anirvan Dutta, Hamid Sadeghian, Ville Kyrki, Sami Haddadin, Gokhan Alcan, Eckehard Steinbach
Organizations: Technical University of Munich (TUM), Germany · Tampere University, Finland · Aalto University, Finland · Mohamed bin Zayed University of Artificial Intelligence (MBZUAI), UAE
Human--robot co-transportation of deformable objects requires predicting object deformation during motion, since obstacle clearance depends on both the grasp points and the unactuated interior. We present a hierarchical planning framework that combines a learned cloth model with a reduced-order whole-body model of a dual-arm mobile manipulator. A physics-residual conditional recurrent variational autoencoder (p-cRVAE) predicts the full cloth configuration from grasp-point observations by learning a residual correction to a computationally efficient linearized physics model, limiting error accumulation over 40-step planning horizon. The predicted cloth dynamics are embedded in a model predictive path integral (MPPI) planner using a reduced-order representation of a dual-arm mobile manipulator that preserves the non-holonomic base constraint and arm workspace limits. An MPC layer subsequently refines the sampled motion into smooth, executable references for whole-body control. The reduced-order formulation achieves tracking performance comparable to the full 17-DoF model while reducing computation time by approximately 80%. Across four co-transportation scenarios and two carrying speeds, the proposed framework maintains cloth-obstacle clearance where a corner-following baseline results in collisions, while whole-body refinement reduces final cloth deformation from 0.93,m to 0.28,m.
Figures & tables
Fig. 1 : Setup for human–robot co-transportation of a rectangular cloth in the presence of obstacles. The planner jointly uses a high-fidelity cloth model (p-cRVAE) with a reduced-order robot model (ROM) to generate feasible hand and base trajectories while following the user and avoiding cloth–obstacle collisions.
Fig. 2 : Structure of the proposed p-cRVAE. The coarse physics Φc advances the cloth state under the commanded grasp motion, and the residual decoder Dθ predicts a per-particle correction conditioned on the belief state hk and the latent variable zk .
Fig. 3 : The performance of the MPC refinement layer is evaluated under the full model vs. The reduced-order model as prediction models for the planner. The full model considers the complete kinematic chains of both arms, while the reduced-order model approximates each arm with a simple link that captures only the corresponding end-effector position and not the arms configurations. The reference trajectory gk:k+Hr generated by the MPPI planner contains jerky motion. Both planners track the reference trajectory and provide a smoother reference to the whole-body controller.
Fig. 4 : Comparison of the planner performance between full and the reduced-order types. Both versions effectively track the reference gkr while smoothing its jerky motion. However, the optimizer using the full model has a 5x higher computational cost than the one using the reduced model.
Fig. 5 : Training data and open-loop prediction accuracy. ( 5(a) ) Grasp-point coordination patterns used to generate the dataset, shown as the settled and final sheet for a representative episode of each pattern. ( 5(b) ) Per-node prediction error over the rollout, with and without the coarse backbone (mean ± one standard deviation over held-out episodes), and spatial error maps over the mesh on a shared scale. Teal markers denote the measured points.
Fig. 6 : Comparison of the MuJoCo simulations and corresponding real-robot experiments for the four human–robot co-transportation scenarios. In real experiment, the green shadow represents the current position of the cloth and red shadow shows the goal position.
Scene , vh
∑canc [ 106 ] ↓
dev T [cm] ↓
∑cact [ 103 ] ↓
clr [cm] ↑
B-I
B-II
Prop.
B-I
B-II
Prop.
B-I
B-II
Prop.
B-I
B-II
Prop.
S1, 0.075
0.10 ± 0.02
1.65 ± 0.03
0.13 ± 0.01
32.5 ± 1.4
50.0 ± 0.2
28.7 ± 4.2
2.00 ± 0.05
2.07 ± 0.14
2.03 ± 0.02
–
–
–
S1, 0.125
0.14 ± 0.01
1.70 ± 0.01
0.16 ± 0.00
33.1 ± 0.5
49.8 ± 0.3
32.2 ± 0.3
2.06 ± 0.05
2.03 ± 0.04
2.02 ± 0.03
–
–
–
S2, 0.075
0.27 ± 0.01
1.93 ± 0.02
0.24 ± 0.01
29.5 ± 2.0
67.6 ± 1.9
29.5 ± 2.3
2.06 ± 0.02
2.50 ± 0.12
2.11 ± 0.17
–
–
–
S2, 0.125
0.53 ± 0.02
2.03 ± 0.03
0.29 ± 0.01
26.1 ± 1.8
66.7 ± 0.3
29.3 ± 0.8
2.03 ± 0.04
2.39 ± 0.11
2.11 ± 0.05
–
–
–
S3, 0.075
0.39 ± 0.01
1.23 ± 0.01
0.25 ± 0.01
22.3 ± 0.3
44.6 ± 0.5
30.1 ± 1.2
2.24 ± 0.03
2.45 ± 0.09
2.26 ± 0.07
–
–
–
TABLE I : End-to-end results on the four evaluation scenes at two carry speeds. Mean ± std over three repeats of a fixed 300-frame run. ∑canchor and ∑cact are integrated over the run. dev T is the midpoint-anchored mean node deviation of the full 225-node sheet at the final frame. Scenes 1–3 carry no obstacle, so clearance is undefined there. B-I: corner-following only, no shape prediction. B-II: no whole-body NMPC.
Ablation
∑canc [ 106 ] ↓
dev T [cm] ↓
∑cact [ 103 ] ↓
clr [cm] ↑
solve [ms] ↓
H=10
0.51 ± 0.00
29.4 ± 0.9
1.11 ± 0.03
− 3.7 ± 0.4
31.9 ± 0.2
H=40 (ours)
0.53 ± 0.01
28.3 ± 1.9
2.18 ± 0.04
1.3 ± 0.3
96.4 ± 0.3
H=60
0.65 ± 0.04
25.3 ± 4.4
2.66 ± 0.09
1.6 ± 0.4
139.3 ± 0.6
wobs×10
0.66 ± 0.02
25.0 ± 2.3
2.15 ± 0.02
9.7 ± 0.2
96.6 ± 1.3
wshape×10
5.92 ± 0.30
22.6 ± 3.0
2.13 ± 0.06
− 7.3 ± 0.6
97.1 ± 0.3
TABLE II : Ablations on scene4 yobs40 at 0.125 m/s, the cell with the tightest time budget. Mean ± std over three repeats. H=40 is the proposed horizon and is read from the same runs as Prop. in Table I .
Manipulating deformable objects (DOs) is challenging due to their high-dimensional state space, underactuated dynamics, and partial observability. In this paper, we propose cRVAE, a lightweight conditional recurrent variational autoencoder that estimates the full DO state from only partial corner-node observations during inference. The resulting model is used as the forward model in a receding-horizon optimal control framework for obstacle-aware collaborative DO manipulation. In simulation on rope and fabric, cRVAE estimates the full DO state from the available corner-node measurements alone, matching the accuracy of a parameter-identified XPBD model. At inference it uses no physical parameters as model inputs and performs no online parameter identification. It also runs approximately 350 times faster on the rope and over 1500 times faster on the fabric per forward pass, keeping horizon-based planning within the 100 ms control budget where XPBD exceeds it already at short horizons. Full-shape estimation from corner sensing at in-loop speed is what makes the model deployable on hardware, which we demonstrate on a Unitree Go2 robot.
Kosar Behnia, Ville Kyrki, Gokhan Alcan
Faculty of Engineering and Natural Sciences, Tampere Univ., Finland. · Electrical Engineering and Automation Department, Aalto Univ., Finland.
We focus on human-robot collaborative transport, a challenging task of broad relevance spanning logistics, manufacturing, and the home, in which a user and a robot work together to relocate a large or heavy object. To act as an effective partner, the robot should reduce the user's effort by contributing to efficient relocation of the object while remaining physically responsive to them. Prior work often addresses these capabilities separately, producing robots that may move the object efficiently but resist user input, or accommodate the user but depend on continuous guidance. Our key insight is that obstacle-constrained collaborative transport requires integrating predictions of human collaborative behavior with compliant robot control. To this end, we introduce PROACT, a framework for human-robot collaborative transport that incorporates anticipation into compliant whole-body control through a learned model of human collaborative behavior. Trained on a large-scale, real-world dataset of dyadic human transport demonstrations, our transformer architecture distills collaborative behavior into predictions of future object motion. Across 108 real-world trials with a 9-DoF mobile manipulator, PROACT reduces mean interaction work by 59.2% and 20.4%, and mean completion time by 12.9% and 6.9%, relative to compliance-only and MPC baselines, respectively. Footage from our experiments can be found at https://youtu.be/qAGvQfVPjbk.
Elvin Yang, Christoforos Mavrogiannis
Department of Robotics, University of Michigan, Ann Arbor, United States
Simulator-in-the-loop optimization offers a promising inference-time mechanism for robot manipulation. It uses a physical simulator as a backend rollout engine to evaluate candidate trajectories in parallel and refine nominal actions online, a paradigm shown to be effective in rigid-body manipulation where state and contact are relatively tractable. We bring this paradigm to real-world cloth manipulation from a single RGB input through three pillars. (i) We design a scalable synthetic-data generation and inference-time rollout pipeline built on FLASH, a deformable-object simulator that provides a practical balance among physical fidelity, numerical stability, and rollout efficiency. (ii) We develop a real-to-sim module, trained purely on synthetic data, that maps a single RGB observation to simulation-compatible cloth state by fusing pretrained visual features with learnable canonical tokens. (iii) We perform online planning by coupling a sparse-mesh rollout backend with prior-guided MPPI, anchored at an offline-distilled policy trajectory, preserving manipulation-relevant deformation and contact while enabling sufficient parallel rollout batches. Real-robot experiments show higher success rates than baseline methods and closed-loop correction under mid-fold perturbations. Project page: https://silr-cloth.github.io/
Xin Liu, Yulin Li, Ziming Li +7
National University of Singapore · Shanghai Jiao Tong University