A pretrained robot foundation policy may execute most of a long-horizon task yet repeatedly fail at a few critical subtasks. Collecting additional full-task demonstrations for supervised fine-tuning (SFT) requires operators to repeat behaviors the policy already performs well. Reinforcement learning (RL) fine-tuning offers a promising path to bridge this gap, but existing approaches struggle to solve long-horizon tasks using only sparse rewards. We present PARTS (Policy Adaptation with RL on Targeted Subtasks), a real-world subtask RL framework that concentrates practice at these bottlenecks while allowing training rollouts to proceed with minimal human intervention. The frozen pretrained policy supplies nominal actions throughout execution, while agent-generated selectors and success verifiers activate residual corrections and provide local outcome rewards. These rewards support learning from successful subtasks even when complete-task successes are scarce. Training combines online RL with success-reweighted retraining, and each retrained residual policy is redeployed to collect further experience. Humans identify bottlenecks during setup and perform physical resets when needed. On bimanual YAM and single-arm Franka tasks, PARTS improves complete-task success from 32% to 61% and from 50% to 95%, respectively, using tens of minutes of real-world RL rollouts per task on average. Compared with existing real-world RL fine-tuning methods, PARTS raises full-task success by more than 25% under the same robot-rollout budget while requiring less human involvement.
Figures & tables
Fig. 2: PARTS architecture. Left: Humans identify bottlenecks from base-policy failures and specify contracts. Coding agents implement selection, verification, and reset programs. Right: Residual policies correct the frozen base policy. Local rewards and per-bottleneck replay support online TD3+BC and retraining on all successes plus a fraction ρ of failures. Retrained policies resume online learning. At evaluation, PARTS enables the switch between the base and residual policies to complete the full task.
Method
Rollout
Reward
Reset
DSRL [ 9 ]
✓
✗
✗
EXPO-FT [ 10 ]
✗
✓ †
✗
RLT [ 7 ]
✗ ∗
✗
✗
PARTS (ours)
✓
✓ ‡
✓ ‡
TABLE I: Human involvement during RL data collection, as published for the baselines and PARTS. ✓: autonomous; ✗: requires a human. Rollout : corrective interventions or handoff decisions during RL rollouts; Reward : success labeling; Reset : scene restoration.
Fig. 3: Tasks on bimanual YAM and single-arm Franka FR3. Colored borders mark states produced by a targeted residual policy, in the color of its subtask name in the row header. Top: earbud insertion, where inserting earbuds requires high precision. Middle: LEGO sorting, where the grasp of each small brick is the recurring bottleneck because the brick size is out-of-distribution for the base policy. Bottom: cable insertion requires precise manipulation.
Fig. 4: PARTS significantly raises success across tasks with different difficulty levels and outperforms RL fine-tuning baselines under a matched robot-rollout budget. Top: full-task evaluation over 20 episodes per method; for each method the light bar is the progress score and the dark bar is binary whole-task success. Earbud progress credits one third per completed stage and success requires all three; cable progress credits one half each for unplugging and insertion and success requires both. LEGO is a repeated pick-and-place task over ten bricks, so we report only its progress score, the fraction of bricks sorted before the deadline, which is more informative than a binary outcome over the whole task. The ablation is reported for binary success only. Bottom: per-stage success measured within the same episodes, for the three earbud bottlenecks and the two cable stages. Hatched bars are PARTS without success-reweighted retraining.
Fig. 5: Robot platforms. (a) Franka FR3, shown with the two routers of the cable task, observed by a fixed ZED camera and a wrist-mounted ZED camera. (b) YAM with I2RT FlexPoint adaptive grippers and three Intel RealSense D405 cameras: a head camera on the overhead pole and one camera on each wrist. Dashed ellipses mark the cameras.
Vision-Language-Action (VLA) policies offer strong general-purpose manipulation priors, but often fail on tight-tolerance, contact-rich assembly due to long-horizon credit assignment and subtask coupling: a state that is geometrically successful for the current skill can be brittle for downstream skills. We show this failure mode in residual reinforcement learning (RL) over a frozen VLA base policy: constant sparse success rewards improve each subtask in isolation yet yield little or no gain when skills are chained, because terminal state quality is uncontrolled. We propose Foresight Residual RL, which optimizes handoff quality by augmenting each subtask's sparse success reward with an offline-estimated foresight value -- the probability of future subtask success conditioned on the terminal state of the current subtask. Concretely, we (i) train a visual foresight predictor from images of terminal states of the base policy, labeled using downstream rollout statistics, and (ii) train residual policies via backward foresight induction, using the predictor output as a reward multiplier. On a three-phase wrench-based nut-tightening assembly task in Isaac Gym (grasp, move-insert, rotate), our method achieves 85.6% full-task success, outperforming standard subtask residual RL (54.5%) and VLA baselines, while leaving per-subtask success unchanged. These results highlight that improving long-horizon performance requires shaping which successful states are produced at each sub-task, not only whether success occurs.
Yuhan Liu, Xinyu Zhang, Litao Liu +1
Department of Computer Science, Rutgers University
Pretrained imitation policies have become a strong foundation for robot manipulation, but they often require online improvement to overcome execution errors, limited dataset coverage, and deployment mismatch. A central question is therefore how reinforcement learning (RL) should adapt policies after offline pretraining. Existing lightweight methods commonly apply residual corrections directly in action space, but this often leads to noisy and poorly structured exploration. In this work, we propose Z-Perturbation Reinforcement Learning (ZPRL), an approach that steers pretrained policies through a compact bottleneck latent rather than through policy weights or output actions. During offline training, we augment the policy with a plug-and-play variational information bottleneck (VIB) module to extract a task-relevant latent interface from observation embeddings. During online finetuning, the base policy is frozen and RL learns only a residual perturbation on this latent, whose decoded representation conditions the frozen action generator. We instantiate ZPRL on flow-matching policies and evaluate it on eight simulation tasks and four real-world tasks. Across diverse manipulation settings, ZPRL improves both sample efficiency and final performance over strong post-training baselines. In the real world, ZPRL improves the average success rate on four tasks by 33.7% over imitation base policies while producing smoother exploration behaviors than an action residual counterpart. These results suggest that a compact, task-aligned bottleneck latent provides an effective interface for online RL adaptation. More videos can be found at https://manutdmoon.github.io/ZPRL/.
Dongjie Yu, Kun Lei, Zhennan Jiang +2
School of Computing and Data Science, The University of Hong Kong, Hong Kong SAR. · Shanghai Qizhi Institute, Shanghai, China. · Shanghai Jiao Tong University, Shanghai, China. +2
Generalist robot policies carry broad manipulation priors from large-scale data, but specializing them to a new task remains the deployment bottleneck. This requires eliciting task-specific behavior from limited demonstrations without degrading their broad capabilities. We introduce Proxy Policy Steering (PPS), an inference-time adaptation method that resolves this challenge by training two lightweight proxy policies whose calibrated velocity-space difference steers the frozen base sampler. A reference proxy models the frozen base's behavior on target-task observations, and a task proxy, initialized from the reference, captures how this behavior changes under task supervision. Their difference forms a calibrated velocity-space residual that steers the frozen base sampler at every denoising step. We identify the conditions under which this residual isolates the change induced by task supervision, and validate them empirically. Because the base is never directly modified, its broad capabilities remain available at inference, including behaviors such as recovery from failure that the demonstrations themselves do not exercise. Adaptation requires only forward velocity predictions from the base, making PPS lightweight to train and applicable even without access to the base's parameters. On 8 real-world and 4 simulation manipulation tasks, PPS lifts the state-of-the-art pi 0.5 base policy by 53% absolute success rate on average, with zero-to-one gains on tasks the base never solves, while preserving the base's broad capabilities. PPS outperforms LoRA fine-tuning, from-scratch specialists, residual policies, and prior inference-time steering methods.