Robot policies deployed in real environments naturally accumulate mixed-quality experience, including successful executions, partial progress, and failures. Although these rollouts provide valuable information for further learning, directly incorporating them into imitation learning may reinforce undesirable behaviors, while offline reinforcement learning often suffers from unreliable value estimation under sparse rewards and limited data coverage. We consider a practical post-deployment setting where learning relies only on naturally accumulated autonomous rollouts, without additional human corrections or exploratory interaction. To effectively exploit such experience, we propose Predictive Action Chunk Learning (PACL). PACL first learns a predictive chunk-level critic that evaluates temporally extended action sequences and augments temporal difference learning with future latent prediction, providing richer supervision for long-horizon value estimation. The learned critic then converts chunk-level Q-values into discrete quality conditions, which guide a diffusion actor to learn jointly from these mixed-quality experiences without treating all behaviors as equivalent supervision. At inference, the actor generates multiple action chunks and the critic selects the highest valued candidate. Experiments across simulated and real-world robot manipulation tasks show that PACL consistently improves the pretrained policy and outperforms strong imitation learning and offline reinforcement learning baselines.
Figures & tables
Fig. 1 : Overview of Predictive Action Chunk Learning (PACL). PACL improves a pretrained policy from human demonstrations and mixed-quality deployment rollouts using a Q-conditioned diffusion actor and a predictive chunk-level critic. The critic evaluates temporally extended actions through TD learning augmented with future latent prediction.
Fig. 2 : Visualization of critic-based action chunk ranking. The predictive critic evaluates multiple candidate action chunks and selects the highest-valued one for execution on Transport and Square tasks.
Fig. 3 : Analysis of critic estimates on Square. Left: state values and chunk-level Q-values are strongly correlated for both successful and failed rollouts. Right: Q-value distributions show clear separation between successful and failed rollouts.
Fig. 4 : Experiment setup. This includes 4 simulation tasks on the Robomimic benchmark and 3 real-robot tasks on the Franka arm.
Method
Data source
Can
Transport
Square
ToolHang
PickCup
StackCup
MoveSpoon
Baseline
DP [ 1 ]
Dh
88.0
84.0
78.8
46.4
84.0
72.0
64.0
IL
SUB [ 7 ]
Dh+Ds+Df
93.6
88.0
80.8
70.0
92.0
84.0
68.0
Self-Imitation
Dh+Ds
94.8
89.2
86.8
78.0
100.0
92.0
72.0
SSDF [ 9 ]
Dh+Ds+Df+
98.4
92.0
88.4
81.2
100.0
96.0
80.0
Offline RL
DQL [ 26 ]
Dh+Ds+Df
92.0
86.4
81.6
56.4
88.0
76.0
72.0
IDQL [ 27 ]
Dh+Ds+Df
91.6
90.0
83.6
58.8
100.0
92.0
76.0
TABLE I : Post-deployment improvement across 4 simulation and 3 real-world manipulation tasks. All methods are warm-started from the pretrained DP and evaluated over 250 simulation rollouts or 25 real-world trials per task.
Fig. 5 : Component ablation of PACL. Both the Q-conditioned actor and predictive action-chunk critic consistently improve performance across all simulation tasks.
Method
N=4
N=8
N=12
N=16
One-step critic
79.2
75.2
63.6
58.8
Predictive one-step critic
85.2
82.8
83.2
76.4
Chunk critic
85.2
73.2
65.2
63.6
PACL
89.2
83.2
86.8
80.4
TABLE II : Ablation of critic design on Transport.
Fig. 6 : Comparison of conditioning strategies for actor post-training with single-sample inference ( N=1 ).
Setting
Data/Epochs
Square
Transport
Baseline DP
200 / 0
78.8
84.0
Post-trained DP
200 / 50
84.0
86.4
PACL-success
300 / 50
77.6
83.2
PACL-rollout
500 / 50
87.6
86.8
PACL-small
350 / 50
88.0
88.8
PACL-full
700 / 50
93.2
96.0
TABLE III : Ablation of post-training data composition and scale. “Data” denotes the total number of trajectories used for training. N=12 candidate chunks are used at inference for all PACL variants.
School of Computing and Data Science, The University of Hong Kong, Hong Kong SAR. · Shanghai Qizhi Institute, Shanghai, China. · Shanghai Jiao Tong University, Shanghai, China. +2