Many Ways to Succeed: Diversity-Driven RL Fine-Tuning for VLA Generalization
Organizations: Nanjing University · Harbin Institute of Technology (Shenzhen) · Australian National University · University of Technology Sydney
Abstract
Reinforcement learning (RL) fine-tuning improves vision-language-action (VLA) policies through closed-loop experience, yet generalization beyond the fine-tuning distribution remains limited. Our analysis reveals a selective reshaping of exploration: RL contracts behavior globally, yet diversifies successful trajectories, elicits success with fewer rollouts, and covers more of the latent task-valid solution space than supervised fine-tuning. Broader successful-mode coverage may provide alternative strategies under distribution shifts. Inspired by this, we introduce DRIVE (Diversity-driven RL fIne-tuning for VLA gEneralization), which turns successful-behavior diversity into an explicit RL objective. DRIVE groups rollouts under matched task conditions, compares their trajectories with temporal alignment, and derives a success-conditioned intrinsic reward from relative behavioral diversity. This design encourages broader coverage of feasible solutions without rewarding diverse failures or superficial timing differences. Across LIBERO-Plus, ManiSkill3, and RoboTwin 2.0, DRIVE improves the average out-of-domain (OOD) performance over vanilla RL fine-tuning by 5.3 points on and 2.0 points on . On a dual-arm AgileX PiPER-X platform, DRIVE further increases average OOD success from 64.1% to 73.3% (+9.2 points), demonstrating gains that persist under physical deployment.
Figures & tables
| Method | LIBERO-Plus Spatial | LIBERO-Plus Object | LIBERO-Plus Goal | LIBERO-Plus Long | ManiSkill3 | RoboTwin 2.0 | Avg. | |||||||
| IND | OOD | IND | OOD | IND | OOD | IND | OOD | IND | OOD | IND | OOD | IND | OOD | |
| SFT | 22.4 | 14.8 | 42.0 | 27.6 | 18.4 | 14.1 | 30.0 | 24.5 | 39.7 | 21.1 | 42.0 | 46.1 | 36.6 | 29.2 |
| Vanilla RFT | 78.8 | 48.9 | 70.8 | 53.4 | 55.6 | 37.9 | 66.0 | 46.1 | 66.2 | 35.5 | 65.9 | 62.1 | 66.6 | 48.1 |
| DRIVE | 79.6 | 50.2 | 74.4 | 53.5 | 54.8 | 39.8 | 66.4 | 50.8 | 68.4 | 40.3 | 70.6 | 71.4 | 69.3 | 53.4 |
Appendix figures & tables8 assets
Supplementary material from the paper’s appendix.
Appendix
| Suite | Clean | Confounding | Target Pose | Scene Theme | Surface | Total |
| Spatial | 10 | 78 | 66 | 44 | 52 | 250 |
| Object | 10 | 77 | 72 | 37 | 54 | 250 |
| Goal | 10 | 81 | 63 | 42 | 54 | 250 |
| Long | 10 | 65 | 50 | 52 | 73 | 250 |
| Perturbation family | Spatial | Object | Goal | Long |
| Objects Layout | 385 | 403 | 425 | 312 |
| Camera Viewpoints | 376 | 396 | 408 | 419 |
| Robot Initial States | 350 | 398 | 409 | 393 |
| Language Instructions | 390 | 354 | 410 | 383 |
| Light Conditions | 292 | 297 | 279 | 274 |
| Background Textures | 258 | 248 | 281 | 289 |
| Split / Category | Evaluation focus | Settings | Episodes |
| IND | Training objects, scenes, and state distribution | 1 | 320 |
| Visual-Language OOD | Visual appearance and language variation | 6 | 1,920 |
| Semantic OOD | Object/receptacle grounding under semantic ambiguity | 4 | 1,280 |
| Execution OOD | Initial-state shifts and execution-time changes | 3 | 960 |
| Total | IND + three OOD categories | 14 | 4,480 |
| Perturbation family | IND | OOD | Description |
| Clean | 50 | 0 | Nominal task configurations |
| Objects | 134 | 34 | Object additions and pose variations |
| Background | 750 | 150 | Scene/background appearance variations |
| Lighting | 0 | 50 | Held-out illumination conditions |
| Camera | 0 | 50 | Held-out camera configurations |
| Robot State | 66 | 16 | Robot initial-state variations |
| Parameter | Spatial | Object | Goal | Long | Spatial | Object | Goal | Long |
| Runner epochs | 200 | 300 | 400 | 200 | 300 | 250 | 100 | 300 |
| Global batch size | 2048 | 2048 | 2048 | 2048 | 2048 | 2048 | 2048 | 2048 |
| PPO update epochs | 4 | 4 | 4 | 4 | 3 | 3 | 3 | 4 |
| Actor learning rate | ||||||||
| Critic learning rate | ||||||||
| ManiSkill3 | Click Bell | Press Stapler | ||||
| Parameter | ||||||
| Runner epochs | 300 | 200 | 100 | 100 | 100 | 50 |
| Global batch size | 5120 | 5120 | 2048 | 2048 | 2048 | 2048 |
| PPO update epochs | 4 | 5 | 5 | 5 | 5 | 5 |
| Actor learning rate | ||||||
| Critic learning rate | ||||||
| Method | Modification |
| Vanilla RFT | Standard PPO + Flow-SDE |
| Higher Noise | Flow-SDE noise |
| KL Regularization | KL penalty |
| Clip-Higher | PPO upper clip |
| DRIVE | Success-conditioned GAK shaping |