Demonstration-conditioned policies provide a natural interface for specifying robot behavior, yet long-horizon manipulation remains difficult when visually similar states recur across different stages or when demonstrations and executions proceed at different speeds. We identify the resulting failure mode as stage confusion and introduce Progress-Aligned Context for Execution (PACE), a stateful method that continually reinterprets a complete demonstration according to realized execution progress. PACE compresses the demonstration into ordered multimodal prompt tokens and uses training-only dual-edge attention supervision to expose its latent stage structure. During execution, an episode-local fast-weight memory causally encodes realized action-observation transitions and modulates prompt cross-attention, producing a progress-aligned context for a unified diffusion action expert without test-time stage labels or stage-specific policies. PACE improves success from 88.9% to 94.0% on LIBERO-Gen Goal Chain, from 79.1% to 83.3% on Spatial Combination, and from 33.3% to 73.3% on the two-step Block Routing tasks. Failure analysis further indicates that structured demonstration alignment and causal execution memory jointly mitigate stage confusion.
Figures & tables
Fig. 1: Stage confusion in long-horizon demonstration-conditioned manipulation. Visually similar observations can recur at different stages while requiring different future actions (top). Observation-only matching may therefore retrieve an earlier stage and repeat completed behavior instead of continuing from the current progress (bottom).
Fig. 2: Architecture overview of PACE. The demonstration prompt encoder produces ordered prompt chunks. During execution, episode-local fast weights encode realized action–observation transitions and modulate observation queries before prompt cross-attention. Training-only dual-edge attention supervision is applied to the demonstration attention. The resulting progress-aligned context and current observation condition a diffusion action expert.
Fig. 3: Stage-structure representation in PACE. (a) Each 20-step multimodal demonstration chunk is aggregated by attention pooling into one ordered prompt token. (b) Detected prompt and target boundaries are monotonically aligned to construct a soft dual-edge target over prompt chunks for attention supervision.
Fig. 5: Complete-task success rates. The horizontal axis denotes the task category, and the vertical axis reports the mean complete-task success rate. (a) Simulation Experiment 1: LIBERO-Gen Spatial Combination and LIBERO-Gen Goal Chain each contain 10 tasks whose stage compositions are unseen during training. Each task is evaluated with 10 rollouts, and the success rate is averaged across tasks. (b) Simulation Experiment 2: Block Routing consists of 2-step and 4-step tasks, with 10 training demonstrations for each task. A 2-step task moves the block to one of the red, green, or blue target regions and then returns it to the home region, yielding three tasks. A 4-step task concatenates an ordered pair of 2-step tasks—for example, Home→Red→Home→Green→Home —yielding 3×3=9 tasks. Each task is evaluated with 10 rollouts, and the mean success rate is reported for each category. (c) Real-robot Block Routing: The task follows the same setup as (b), using 15 training demonstrations and 10 evaluation rollouts per task; the mean success rate is reported for each category.
Fig. 6: Complete outcome composition on the two-step Block Routing task over 30 rollouts. Each horizontal bar partitions 30 episodes into successes, wrong-stage failures, wrong-color failures, and wrong-grasp failures. Failure subsegments are uniformly compressed to retain the common 30-episode endpoint, while labels report exact counts. The corresponding success rate is reported at right.
Fig. 7: Attention maps for two-step Block Routing: (a) BPP and (b) PACE w/o FWM show representative failures, while (c) Full PACE shows a representative success. Rows are aligned to four intended stages— To block , To target tray , To block , and Return block home . Blue intensity represents normalized attention weight, gray dashed lines separate intervals, and red curves mark the attention center of mass.
Fig. 8: Counterfactual history conditioning with the current observation and demonstration held fixed. The upper panel directly contrasts the visually aliased first and second To block occurrences. The lower panel compares matched, earlier, later, and shuffled histories for a randomly selected sample. Black markers denote the attention center of mass, and red lines denote the matched-stage boundaries.
Institute for AI Industry Research (AIR), Tsinghua University · School of Electronic Information Engineering, Beihang University · AIR Wuxi Innovation Center, Tsinghua University +3