Robot policies deployed in real environments naturally accumulate mixed-quality experience, including successful executions, partial progress, and failures. Although these rollouts provide valuable information for further learning, directly incorporating them into imitation learning may reinforce undesirable behaviors, while offline reinforcement learning often suffers from unreliable value estimation under sparse rewards and limited data coverage. We consider a practical post-deployment setting where learning relies only on naturally accumulated autonomous rollouts, without additional human corrections or exploratory interaction. To effectively exploit such experience, we propose Predictive Action Chunk Learning (PACL). PACL first learns a predictive chunk-level critic that evaluates temporally extended action sequences and augments temporal difference learning with future latent prediction, providing richer supervision for long-horizon value estimation. The learned critic then converts chunk-level Q-values into discrete quality conditions, which guide a diffusion actor to learn jointly from these mixed-quality experiences without treating all behaviors as equivalent supervision. At inference, the actor generates multiple action chunks and the critic selects the highest valued candidate. Experiments across simulated and real-world robot manipulation tasks show that PACL consistently improves the pretrained policy and outperforms strong imitation learning and offline reinforcement learning baselines.
Figures & tables
Fig. 1 : Overview of Predictive Action Chunk Learning (PACL). PACL improves a pretrained policy from human demonstrations and mixed-quality deployment rollouts using a Q-conditioned diffusion actor and a predictive chunk-level critic. The critic evaluates temporally extended actions through TD learning augmented with future latent prediction.
Fig. 2 : Visualization of critic-based action chunk ranking. The predictive critic evaluates multiple candidate action chunks and selects the highest-valued one for execution on Transport and Square tasks.
Fig. 3 : Analysis of critic estimates on Square. Left: state values and chunk-level Q-values are strongly correlated for both successful and failed rollouts. Right: Q-value distributions show clear separation between successful and failed rollouts.
Fig. 4 : Experiment setup. This includes 4 simulation tasks on the Robomimic benchmark and 3 real-robot tasks on the Franka arm.
Method
Data source
Can
Transport
Square
ToolHang
PickCup
StackCup
MoveSpoon
Baseline
DP [ 1 ]
Dh
88.0
84.0
78.8
46.4
84.0
72.0
64.0
IL
SUB [ 7 ]
Dh+Ds+Df
93.6
88.0
80.8
70.0
92.0
84.0
68.0
Self-Imitation
Dh+Ds
94.8
89.2
86.8
78.0
100.0
92.0
72.0
SSDF [ 9 ]
Dh+Ds+Df+
98.4
92.0
88.4
81.2
100.0
96.0
80.0
Offline RL
DQL [ 26 ]
Dh+Ds+Df
92.0
86.4
81.6
56.4
88.0
76.0
72.0
IDQL [ 27 ]
Dh+Ds+Df
91.6
90.0
83.6
58.8
100.0
92.0
76.0
TABLE I : Post-deployment improvement across 4 simulation and 3 real-world manipulation tasks. All methods are warm-started from the pretrained DP and evaluated over 250 simulation rollouts or 25 real-world trials per task.
Fig. 5 : Component ablation of PACL. Both the Q-conditioned actor and predictive action-chunk critic consistently improve performance across all simulation tasks.
Method
N=4
N=8
N=12
N=16
One-step critic
79.2
75.2
63.6
58.8
Predictive one-step critic
85.2
82.8
83.2
76.4
Chunk critic
85.2
73.2
65.2
63.6
PACL
89.2
83.2
86.8
80.4
TABLE II : Ablation of critic design on Transport.
Fig. 6 : Comparison of conditioning strategies for actor post-training with single-sample inference ( N=1 ).
Setting
Data/Epochs
Square
Transport
Baseline DP
200 / 0
78.8
84.0
Post-trained DP
200 / 50
84.0
86.4
PACL-success
300 / 50
77.6
83.2
PACL-rollout
500 / 50
87.6
86.8
PACL-small
350 / 50
88.0
88.8
PACL-full
700 / 50
93.2
96.0
TABLE III : Ablation of post-training data composition and scale. “Data” denotes the total number of trajectories used for training. N=12 candidate chunks are used at inference for all PACL variants.
Recent vision-language-action and diffusion-based robot policies often use action chunking, where each policy query predicts a sequence of future actions and the robot executes an open-loop prefix before re-querying. While this interface improves local motion continuity, deployment still requires choosing the execution horizon: how much of each predicted chunk should be executed before acquiring a new observation. However, our experiments show that success is strongly task-dependent and non-monotonic with respect to the execution horizon, making a single constant horizon an unreliable deployment rule. We propose PACE (Phase-Aware Chunk Execution), a training-free test-time execution method that selects the execution horizon online from the predicted chunk itself. PACE exploits the phase-dependent kinematic structure of manipulation trajectories by identifying low-speed transition points in the predicted speed profile and using them as candidate replanning boundaries. Because PACE uses only the predicted action chunk, it is plug-and-play and requires no retraining or access to policy internals. We validate PACE through large-scale evaluations in both simulation and real-robot settings. On 50 RoboTwin2.0 tasks, PACE raises the average success rate from 57.8% to 64.2%. In real-robot experiments on bimanual ALOHA and single-arm Franka platforms, PACE improves the average task score from 60.7 to 77.7 and the average success rate from 50.7% to 70.4%. Ablations and rollout-level analyses show that PACE adapts execution horizons across manipulation phases, shortening near transitions while preserving longer execution during coherent motion.
Junnan Nie, Jiayi Li, Chenghao Liu +5
Peking University · Peking University. · JD Explore Academy. +1
Action chunking provides temporal abstraction in reinforcement learning by selecting short action sequences instead of individual actions, but many existing approaches face two limitations in high-dimensional robotic control. First, many rely on value functions over action chunks, which can be difficult to learn as action dimensionality and chunk length grow. Second, executing chunks open-loop removes within-chunk feedback, limiting reactivity in contact-rich tasks. We present Action Chunking PPO (ACPPO), a PPO extension that uses a chunked actor while retaining a standard state-value critic, thereby avoiding chunked Q-functions. We further propose ACPPO-Corr, which augments the chunk planner with a stepwise feedback corrector that adjusts planned actions online within each chunk. Across 25 simulated robotics tasks from IsaacGym and Bi-DexHands, spanning locomotion, arm manipulation, and dexterous hand-object interaction, ACPPO-Corr achieves the strongest aggregate performance among evaluated methods and performs best on both decision-frequency-sensitive and decision-frequency-neutral task subsets. Ablations show that moderate chunk lengths work best and that corrector regularization is important for balancing chunk-level planning with local feedback. These results suggest that action chunking can be effective in online PPO when chunk-level planning is paired with closed-loop correction. The code is available at: https://github.com/hshhahn/ACPPO.
Pretrained imitation policies have become a strong foundation for robot manipulation, but they often require online improvement to overcome execution errors, limited dataset coverage, and deployment mismatch. A central question is therefore how reinforcement learning (RL) should adapt policies after offline pretraining. Existing lightweight methods commonly apply residual corrections directly in action space, but this often leads to noisy and poorly structured exploration. In this work, we propose Z-Perturbation Reinforcement Learning (ZPRL), an approach that steers pretrained policies through a compact bottleneck latent rather than through policy weights or output actions. During offline training, we augment the policy with a plug-and-play variational information bottleneck (VIB) module to extract a task-relevant latent interface from observation embeddings. During online finetuning, the base policy is frozen and RL learns only a residual perturbation on this latent, whose decoded representation conditions the frozen action generator. We instantiate ZPRL on flow-matching policies and evaluate it on eight simulation tasks and four real-world tasks. Across diverse manipulation settings, ZPRL improves both sample efficiency and final performance over strong post-training baselines. In the real world, ZPRL improves the average success rate on four tasks by 33.7% over imitation base policies while producing smoother exploration behaviors than an action residual counterpart. These results suggest that a compact, task-aligned bottleneck latent provides an effective interface for online RL adaptation. More videos can be found at https://manutdmoon.github.io/ZPRL/.
Dongjie Yu, Kun Lei, Zhennan Jiang +2
School of Computing and Data Science, The University of Hong Kong, Hong Kong SAR. · Shanghai Qizhi Institute, Shanghai, China. · Shanghai Jiao Tong University, Shanghai, China. +2