VAMPS: Visual and Motor Policies from Sampling-Based Planning
Authors: Mohamed Yassine Kabouri, Pietro Noah Crestaz, Quang-Nam Nguyen, Qilong Cheng, Ludovic Righetti, Nicolas Mansard
Organizations: LAAS-CNRS, Université de Toulouse, CNRS, Toulouse, France. · Machines in Motion Laboratory, New York University, New York, USA. · Industrial Engineering Department, University of Trento, Trento, Italy. · Artificial and Natural Intelligence Toulouse Institute (ANITI), Toulouse.
Learning robot policies directly on physical systems remains difficult because data collection is costly and policy exploration can be unsafe. We introduce Visual and Motor Policies from Sampling-Based Planning (VAMPS), a framework that uses Model Predictive Path Integral (MPPI) control to train reusable policies without human demonstrations. VAMPS supports two training modes. For one-step proprioceptive policies, it operates iteratively in simulation: the policy warm-starts MPPI, and the refined trajectories provide new supervision as the policy changes. A learned terminal value improves short-horizon planning, while an Implicit Q-Learning (IQL) critic guides the policy update. Iterative refinement outperforms training once on frozen MPPI data, and we transfer the learned locomotion policy to a Unitree Go2. For visuomotor policies, VAMPS operates directly from real-robot data. MPPI uses task-specific state estimates to plan and execute trajectories while recording RGB and sensor observations on a Flexiv Rizon 10S. Action Chunking with Transformers predicts action chunks, reducing the effective prediction horizon, and is trained offline on this fixed dataset. We demonstrate visuomotor pick-and-place and force-aware whiteboard erasing. In the latter task, the policy additionally observes the measured 6-D wrench and desired normal force. These results show that VAMPS can learn policies either in simulation followed by hardware transfer or directly from autonomously collected real-robot data.
Figures & tables
Fig. 1 : Overview of the two VAMPS modes. (a) VAMPS-Iterative: a one-step proprioceptive policy warm-starts MPPI, which refines its behavior and provides new supervision. A terminal value improves short-horizon planning, and an IQL critic weights the policy updates. (b) VAMPS-Batch: MPPI uses task-specific estimates to autonomously collect RGB, sensor and action episodes on the real robot, and ACT trains offline on the resulting fixed dataset.
Fig. 2 : The hardware that we consider, four frames each in temporal order. Top: bounding policy of the Unitree Go2 under a forward velocity command, with a proprioceptive policy trained in simulation. Middle: pick-and-place on the Flexiv Rizon 10S. Bottom: whiteboard erasing on the same arm. The two arm tasks are visuomotor and are learned directly from hardware data. Videos are provided in the supplementary video.
Fig. 3 : The three VAMPS-Iterative coupling strategies. Each task shows a Gaussian policy and a deterministic policy. Shaded bands: 95 % CI over 4 seeds.
Fig. 4 : Left: median primal residual over four seeds: BC already follows the planner to 0.24 with no consensus enforced. Right: 95th percentile of ∣yt,j(i)∣ over states and action coordinates, on one Go2 seed. Colorbar in action units; 1 is the bound of πθ .
Fig. 5 : Test return on Panda pick-and-place. Left: planning horizons T∈{16,8,4} with and without the learned terminal value (TV). Right: VAMPS-Iterative without a critic and with the critic for different AWR temperatures β ; both retain the terminal value. Shaded bands: 95 % CI across seeds.
Fig. 6 : Sensitivity to the hyperparameters. Left: penalty coefficient ρ , with the batch size fixed at 1024 . Right: batch size, with ρ fixed at 0.005 . Shaded bands: 95 % CI over 4 seeds.
Method
Episode length
Return
Success
PPO
80
279±84
0.00
PPO
500
975±10
0.98±0.03
PPO
1000
894±163
0.74±0.49
BC (frozen)
–
290±85
0.02±0.03
IQL (frozen)
–
240±40
0.01±0.01
VAMPS-Iterative
80
1109±76
0.98±0.02
TABLE I : Baselines on Panda: mean ± 95 % CI over four seeds at each method’s best checkpoint. PPO is trained at the listed episode length and evaluated over 80 steps, as VAMPS-Iterative.
Fig. 7 : Desired and measured normal force during one successful VAMPS-Batch erasing trial shown in Fig. 2 .
Department of Mechanical Engineering, Yale University, New Haven, CT 06520, USA · Department of Computer Science, Yale University, New Haven, CT 06520, USA · School of Electrical and Computer Engineering, University of Sydney, Sydney, NSW 2006, Australia
School of Information, Renmin University of China, Beijing, China · XYZ Embodied AI, Beijing, China · Key Laboratory of Data Engineering and Knowledge Engineering, Beijing, China +3