cs.LGSep 24, 2026

Learning from Mixed-Quality Deployment Experience for Robot Manipulation

Authors: Yangang Ren, Yujie Yan, Zirui Li, Jiaming Guo, Di Zeng, Ji Tao, Lan Yu, Xuesong Tian, +1 more

Organizations: School of Mechanical and Aerospace Engineering, Nanyang Technological University, Singapore 639798 · Chongqing Changan Automobile Co., Ltd, Chongqing 400023, China · Guangzhou Cloudbutterfly Technology Co., Ltd., Guangzhou 510220, China

Abstract

Robot policies deployed in real environments naturally accumulate mixed-quality experience, including successful executions, partial progress, and failures. Although these rollouts provide valuable information for further learning, directly incorporating them into imitation learning may reinforce undesirable behaviors, while offline reinforcement learning often suffers from unreliable value estimation under sparse rewards and limited data coverage. We consider a practical post-deployment setting where learning relies only on naturally accumulated autonomous rollouts, without additional human corrections or exploratory interaction. To effectively exploit such experience, we propose Predictive Action Chunk Learning (PACL). PACL first learns a predictive chunk-level critic that evaluates temporally extended action sequences and augments temporal difference learning with future latent prediction, providing richer supervision for long-horizon value estimation. The learned critic then converts chunk-level Q-values into discrete quality conditions, which guide a diffusion actor to learn jointly from these mixed-quality experiences without treating all behaviors as equivalent supervision. At inference, the actor generates multiple action chunks and the critic selects the highest valued candidate. Experiments across simulated and real-world robot manipulation tasks show that PACL consistently improves the pretrained policy and outperforms strong imitation learning and offline reinforcement learning baselines.

Figures & tables

Explore similar work

CardsList
  1. PACE: Phase-Aware Chunk Execution for Robot Policies with Action Chunking

    May 30, 2026Junnan Nie, Jiayi Li, Chenghao Liu +5Action ChunksRobot Policies

  2. Action Chunking Proximal Policy Optimization with Feedback Correction

    Sep 28, 2026Sanghyun Hahn, Jonghyun ChoiAction ChunksProximal Policy Optimization

  3. Beyond Action Residuals: Real-World Robot Policy Steering via Bottleneck Latent Reinforcement Learning

    May 19, 2026Dongjie Yu, Kun Lei, Zhennan Jiang +2Robot PoliciesTemporally Coherent Imitation Learning