Demonstration-conditioned policies provide a natural interface for specifying robot behavior, yet long-horizon manipulation remains difficult when visually similar states recur across different stages or when demonstrations and executions proceed at different speeds. We identify the resulting failure mode as stage confusion and introduce Progress-Aligned Context for Execution (PACE), a stateful method that continually reinterprets a complete demonstration according to realized execution progress. PACE compresses the demonstration into ordered multimodal prompt tokens and uses training-only dual-edge attention supervision to expose its latent stage structure. During execution, an episode-local fast-weight memory causally encodes realized action-observation transitions and modulates prompt cross-attention, producing a progress-aligned context for a unified diffusion action expert without test-time stage labels or stage-specific policies. PACE improves success from 88.9% to 94.0% on LIBERO-Gen Goal Chain, from 79.1% to 83.3% on Spatial Combination, and from 33.3% to 73.3% on the two-step Block Routing tasks. Failure analysis further indicates that structured demonstration alignment and causal execution memory jointly mitigate stage confusion.
Figures & tables
Fig. 1: Stage confusion in long-horizon demonstration-conditioned manipulation. Visually similar observations can recur at different stages while requiring different future actions (top). Observation-only matching may therefore retrieve an earlier stage and repeat completed behavior instead of continuing from the current progress (bottom).
Fig. 2: Architecture overview of PACE. The demonstration prompt encoder produces ordered prompt chunks. During execution, episode-local fast weights encode realized action–observation transitions and modulate observation queries before prompt cross-attention. Training-only dual-edge attention supervision is applied to the demonstration attention. The resulting progress-aligned context and current observation condition a diffusion action expert.
Fig. 3: Stage-structure representation in PACE. (a) Each 20-step multimodal demonstration chunk is aggregated by attention pooling into one ordered prompt token. (b) Detected prompt and target boundaries are monotonically aligned to construct a soft dual-edge target over prompt chunks for attention supervision.
Fig. 5: Complete-task success rates. The horizontal axis denotes the task category, and the vertical axis reports the mean complete-task success rate. (a) Simulation Experiment 1: LIBERO-Gen Spatial Combination and LIBERO-Gen Goal Chain each contain 10 tasks whose stage compositions are unseen during training. Each task is evaluated with 10 rollouts, and the success rate is averaged across tasks. (b) Simulation Experiment 2: Block Routing consists of 2-step and 4-step tasks, with 10 training demonstrations for each task. A 2-step task moves the block to one of the red, green, or blue target regions and then returns it to the home region, yielding three tasks. A 4-step task concatenates an ordered pair of 2-step tasks—for example, Home→Red→Home→Green→Home —yielding 3×3=9 tasks. Each task is evaluated with 10 rollouts, and the mean success rate is reported for each category. (c) Real-robot Block Routing: The task follows the same setup as (b), using 15 training demonstrations and 10 evaluation rollouts per task; the mean success rate is reported for each category.
Fig. 6: Complete outcome composition on the two-step Block Routing task over 30 rollouts. Each horizontal bar partitions 30 episodes into successes, wrong-stage failures, wrong-color failures, and wrong-grasp failures. Failure subsegments are uniformly compressed to retain the common 30-episode endpoint, while labels report exact counts. The corresponding success rate is reported at right.
Fig. 7: Attention maps for two-step Block Routing: (a) BPP and (b) PACE w/o FWM show representative failures, while (c) Full PACE shows a representative success. Rows are aligned to four intended stages— To block , To target tray , To block , and Return block home . Blue intensity represents normalized attention weight, gray dashed lines separate intervals, and red curves mark the attention center of mass.
Fig. 8: Counterfactual history conditioning with the current observation and demonstration held fixed. The upper panel directly contrasts the visually aliased first and second To block occurrences. The lower panel compares matched, earlier, later, and shuffled histories for a randomly selected sample. Black markers denote the attention center of mass, and red lines denote the matched-stage boundaries.
Recent vision-language-action and diffusion-based robot policies often use action chunking, where each policy query predicts a sequence of future actions and the robot executes an open-loop prefix before re-querying. While this interface improves local motion continuity, deployment still requires choosing the execution horizon: how much of each predicted chunk should be executed before acquiring a new observation. However, our experiments show that success is strongly task-dependent and non-monotonic with respect to the execution horizon, making a single constant horizon an unreliable deployment rule. We propose PACE (Phase-Aware Chunk Execution), a training-free test-time execution method that selects the execution horizon online from the predicted chunk itself. PACE exploits the phase-dependent kinematic structure of manipulation trajectories by identifying low-speed transition points in the predicted speed profile and using them as candidate replanning boundaries. Because PACE uses only the predicted action chunk, it is plug-and-play and requires no retraining or access to policy internals. We validate PACE through large-scale evaluations in both simulation and real-robot settings. On 50 RoboTwin2.0 tasks, PACE raises the average success rate from 57.8% to 64.2%. In real-robot experiments on bimanual ALOHA and single-arm Franka platforms, PACE improves the average task score from 60.7 to 77.7 and the average success rate from 50.7% to 70.4%. Ablations and rollout-level analyses show that PACE adapts execution horizons across manipulation phases, shortening near transitions while preserving longer execution during coherent motion.
Junnan Nie, Jiayi Li, Chenghao Liu +5
Peking University · Peking University. · JD Explore Academy. +1
An action chunk can span several stages of a manipulation task, yet a label for its first step describes only the current stage. We introduce Chunk-Aligned Semantic Distillation (CASD), which derives semantic targets for entire action chunks. An offline vision--language model segments demonstrations into described stages. Their occupancy within each action chunk determines a weighted semantic target, including transitions between stages. A CASD generator learns to predict this target from the current observation, robot state, and task instruction. We then freeze the generator and train a policy conditioned on its predictions. The semantic branch runs once per policy query, without online VLM calls or reasoning-trace decoding. Teacher matching on annotated LIBERO training episodes is above chance for both single-stage and boundary-crossing chunks. We evaluate three Fast-WAM variants and a DreamZero integration across four benchmarks, including distribution shifts on LIBERO-Plus. Compared with published references, IDM+CASD reaches 98.9% versus 98.0% average success on LIBERO, while Uncond falls below its reference. Joint+CASD reaches 93.0% versus 90.6% on RoboTwin 2.0, and DreamZero+CASD reaches a 47.9% four-category MolmoSpaces manipulation average versus 40.7%. Performance varies across backbone integrations.
Generalist robot policies aim to map multimodal observations and linguistic task instructions to actions across diverse tasks. However, existing methods typically represent the future as a fixed, short video-action chunk. This short-term future captures local scene evolution for action execution, but it does not explicitly describe the stage-level future that specifies how a task should progress from its current stage to the next. We therefore distinguish two complementary futures for robot manipulation: a short-term physical future to capture local scene evolution and a stage-level semantic future to represent task progress. We introduce JEPA-WAM, which augments a Motus-based World Action Model (WAM) with Stage-JEPA, a goal-conditioned Joint-Embedding Predictive Architecture (JEPA) predictor. Given the current observation and task instruction, Stage-JEPA uses a frozen V-JEPA2 encoder to extract the current-state representation and predicts the latent target of the next inferred stage. Across 50 RoboTwin 2.0 tasks in clean and randomized environments, JEPA-WAM achieves 90.25% overall success and reduces the mean number of execution steps in successful rollouts by 5.97% relative to the strongest baseline.
Xiao Liu, Yuguang Yang, Xi Wang +6
Institute for AI Industry Research (AIR), Tsinghua University · School of Electronic Information Engineering, Beihang University · AIR Wuxi Innovation Center, Tsinghua University +3