cs.ROOct 5, 2026

How (and How Not) to Use Data Augmentation in VLA Post-Training

Authors: Bram Grooten, Joaquin Vanschoren

Organizations: TU Eindhoven

Abstract

Vision-language-action (VLA) models currently demonstrate strong performance in a wide range of real-world robotics tasks. However, they often still lack the generalization ability to handle large visual out-of-distribution shifts. Post-training of VLAs with reinforcement learning (RL) has been shown to benefit robustness, but significant room for improvement remains. In this work, we systematically study the effect of image augmentation on VLA post-training. We find that it is crucial to augment only the critic module during RL updates, while leaving the actor's input clean during both rollouts and updates. For π0.5π_{0.5} and GR00T N1.5 this raises out-of-distribution success on LIBERO-Plus by 7.87.8 and 10.010.0 points respectively, while augmenting the actor collapses training entirely. We investigate a range of augmentation types and strengths, and provide practical recommendations for improving generalization in VLA post-training.

Figures & tables

Appendix figures & tables6 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. RL2^2-VLA: Adaptive RL Latent Compositional Steering with Test-Time Scaling for Vision-Language-Action Models

    Jul 29, 2026Derek Ming Siang Tan, Shailesh Shailesh, Srikrishna Iyer +4Inference-Time SteeringVideo Latents

  2. Seeing Realism from Simulation: Efficient Video Transfer for Vision-Language-Action Data Augmentation

    May 4, 2026Chenyu Hui, Xiaodi Huang, Siyu Xu +5Diffusion-Based Vision-Language-ActionsData Augmentation

  3. Scaling by Diversified Experience for Vision-Language-Action Models

    Jun 8, 2026Leiyu Wang, Zhaofengnian Wang, Xueqi Li +3Decoupling