cs.ROSep 30, 2026

Online Evolution Strategy for Flow-Matching VLA Policies via Self-Supervised Trajectory Distribution Optimization

Authors: Gongxin Yao, Yongsheng Zhao, Jiayin Deng, Deng Liang, Han Gao, Lei Zhao, Baoping Cheng

Organizations: China Mobile (Hangzhou) Information Technology Co., Ltd., Hangzhou, 311121, China

Abstract

Vision-Language-Action (VLA) models based on generative frameworks, such as Flow Matching, have recently achieved impressive performance in robotic manipulation. Unlike deterministic policies, Flow Matching enables VLA models to learn conditional action trajectory distributions, where latent noise vectors induce different actions under the same task scenario. However, we observe that these distributions are often ill-formed, with successful and failed behaviors coexisting while considerable probability mass remains in unfavorable regions. To this end, we propose Online-ES, an online adaptation framework for Flow Matching VLAs based on Evolution Strategy (ES), which refines the learned action trajectory distribution through interaction feedback. Instead of pruning the latent noise space, our method performs evolutionary exploration directly in the action trajectory space, where diverse trajectories generated by Flow Matching provide candidate solutions for adaptation. By perturbing sampled trajectories and evaluating their execution outcomes, we derive a self-supervised MSE objective that transfers the evolution direction from trajectory space into model parameter space. Mathematically, we prove that the proposed objective provides an unbiased estimator of the optimal evolution direction. Moreover, we also incorporate failure experiences as negative feedback to regularize the evolution direction, steering the policy away from previously explored failure regions. Experiments in both simulation and real-world environments demonstrate that Online-ES achieves policy improvement comparable to reinforcement fine-tuning, without learning a value model or computing advantages.

Figures & tables

Explore similar work

Oct 11, 2025cs.LG

Reinforcement Fine-Tuning of Flow-Matching Policies for Vision-Language-Action Models

Vision-Language-Action (VLA) models such as OpenVLA, Octo, and π0π_0 have shown strong generalization by leveraging large-scale demonstrations, yet their performance is still fundamentally constrained by the quality and coverage of supervised data. Reinforcement learning (RL) provides a promising path for improving and fine-tuning VLAs through online interaction. However, conventional policy gradient methods are computationally infeasible in the context of flow-matching based models due to the intractability of the importance sampling process, which requires explicit computation of policy ratios. To overcome this limitation, we propose Flow Policy Optimization (FPO) algorithm, which reformulates importance sampling by leveraging per-sample changes in the conditional flow-matching objective. Furthermore, FPO achieves stable and scalable online reinforcement fine-tuning of the π0π_0 model by integrating structure-aware credit assignment to enhance gradient efficiency, clipped surrogate objectives to stabilize optimization, multi-step latent exploration to encourage diverse policy updates, and a Q-ensemble mechanism to provide robust value estimation. We evaluate FPO on the LIBERO benchmark and the ALOHA simulation task against supervised, preference-aligned, diffusion-based, autoregressive online RL, and π0π_0-FAST baselines, observing consistent improvements over the imitation prior and strong alternatives with stable learning under sparse rewards. In addition, ablation studies and analyses of the latent space dynamics further highlight the contributions of individual components within FPO, validating the effectiveness of the proposed computational modules and the stable convergence of the conditional flow-matching objective during online RL.
Jun 3, 2026cs.RO

FlowPRO: Reward-Free Reinforced Fine-Tuning of Flow-Matching VLAs via Proximalized Preference Optimization

Post-training Vision-Language-Action (VLA) models into policies that can be reliably deployed on real robots remains a major bottleneck. SFT and DAgger exploit failure signals only indirectly, and reward-based RL is bottlenecked by the difficulty of real-world reward design and of training reliable critics. We present FlowPRO, a reward-free offline reinforced fine-tuning framework for flow-matching VLAs. Algorithmically, we propose RPRO (Robotic Flow-matching Proximalized Preference Optimization), a preference-optimization objective tailored to the flow-matching action head of VLA models. RPRO pairs a contrastive optimizer with an explicit proximal regularizer that anchors the absolute magnitude of the implicit reward, thereby eliminating the reward-hacking failure mode of plain Flow-DPO. On the data side, a teleoperated intervention-and-rollback paradigm produces naturally paired positive and negative trajectories (τw,τl)(τ^w, τ^l) on a real robot from a single operator action; a Smooth Interpolation procedure, combined with batch mixing, then converts these sparse corrections into dense per-state supervision while preserving the base policy's capabilities. On four long-horizon bimanual tasks, FlowPRO attains the highest success rate, outperforming four representative baselines, and ablations confirm the contribution of each loss component.
Jul 30, 2026cs.RO

RedFlow: Redirect Failure into Action-level Corrections for Flow-matching VLA Policy

Reinforcement learning (RL) can improve Vision-Language-Action (VLA) policies from deployment experience, but reward- and preference-based RL primarily identifies desirable behaviors without specifying how to correct failed actions, underutilizing failure trajectories and limiting sample efficiency. Can such corrections be derived from fixed rollouts? Our key insight is that rollouts with different outcomes may contain action chunks executed in similar states, enabling higher-quality chunks to provide locally supported corrective references. Building on this insight, we introduce \textbf{RedFlow}, an offline post-training method for flow-matching VLA policies. \emph{Execution-Context Matching} groups chunks using a compact representation of estimated task progress and robot proprioception. \emph{Quality-Guided Action Redirection} assigns signed chunk-quality scores and aggregates higher-quality chunks into corrective targets, reinforcing high-quality chunks, suppressing low-quality chunks, and redirecting correctable chunks toward their targets. RedFlow requires neither external HIL corrections nor online data collection during post-training. Across four LIBERO suites, RedFlow improves average success from 56.2% to 68.2%, outperforming the strongest evaluated offline baseline, AWR (62.3%), by 5.9 points. Across three real-robot tasks, it improves average success from 56.7% to 74.7%. On LIBERO-Spatial, RedFlow reaches 75.8% success with 1{,}536 fixed rollouts, while the evaluated online methods require 8.7--16×\times as many fresh post-training rollouts to reach the same threshold.