cs.ROOct 7, 2026

Many Ways to Succeed: Diversity-Driven RL Fine-Tuning for VLA Generalization

Authors: Haoru Li, Jinmei Liu, Zhiyong Wang, Xiaoming Li, Zhenhong Sun, Daoyi Dong, Chunlin Chen, Zhi Wang

Organizations: Nanjing University · Harbin Institute of Technology (Shenzhen) · Australian National University · University of Technology Sydney

Abstract

Reinforcement learning (RL) fine-tuning improves vision-language-action (VLA) policies through closed-loop experience, yet generalization beyond the fine-tuning distribution remains limited. Our analysis reveals a selective reshaping of exploration: RL contracts behavior globally, yet diversifies successful trajectories, elicits success with fewer rollouts, and covers more of the latent task-valid solution space than supervised fine-tuning. Broader successful-mode coverage may provide alternative strategies under distribution shifts. Inspired by this, we introduce DRIVE (Diversity-driven RL fIne-tuning for VLA gEneralization), which turns successful-behavior diversity into an explicit RL objective. DRIVE groups rollouts under matched task conditions, compares their trajectories with temporal alignment, and derives a success-conditioned intrinsic reward from relative behavioral diversity. This design encourages broader coverage of feasible solutions without rewarding diverse failures or superficial timing differences. Across LIBERO-Plus, ManiSkill3, and RoboTwin 2.0, DRIVE improves the average out-of-domain (OOD) performance over vanilla RL fine-tuning by 5.3 points on π0π_0 and 2.0 points on π0.5π_{0.5}. On a dual-arm AgileX PiPER-X platform, DRIVE further increases average OOD success from 64.1% to 73.3% (+9.2 points), demonstrating gains that persist under physical deployment.

Figures & tables

Appendix figures & tables8 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. RL2^2-VLA: Adaptive RL Latent Compositional Steering with Test-Time Scaling for Vision-Language-Action Models

    Jul 29, 2026Derek Ming Siang Tan, Shailesh Shailesh, Srikrishna Iyer +4Vision-Language-Action ModelsTest-Time Adaptation of VLA Models

  2. What to Ignore, What to React: Visually Robust RL Fine-Tuning of VLA Models

    May 13, 2026Yuanfang Peng, Jingjing Fu, Chuheng Zhang +6VLM RobustnessDomain Generalization

  3. FORCE: Efficient VLA Reinforcement Fine-Tuning via Value-Calibrated Warm-up and Self-Distillation

    Jun 24, 2026Shuyi Zhang, Yunfan Lou, Hongyang Cheng +8RL for RoboticsVision-Language-Action Models