cs.ROOct 7, 2026

Diagnosing and Recovering from Observation-Space Shift at Long-Horizon Skill Seams

Authors: Pranav Wagh, Yu Fang, Yue Yang, Mingyu Ding

Organizations: Department of Computer Science, University of North Carolina at Chapel Hill, NC, USA.

Abstract

Long-horizon robotic manipulation is often built by chaining independently trained skills. Although each skill can be reliable in isolation, performance degrades sharply when skills are chained: each downstream skill must start from the state its predecessor leaves behind rather than from its training distribution. We study this failure mode, Observation-Space Shift (OSS), and ask what causes these skill-seam failures. Using privileged simulator resets, we find that the dominant shift comes from displaced scene state (e.g., an open drawer or secondary objects left behind by earlier skills), not from the robot's joint configuration or the object the downstream skill manipulates. To test this diagnosis, we build a fully learned detect-restore-resume system: a task-progress monitor detects the stall, a learned policy restores the displaced scene components, and seam-robust fine-tuning lets the skill resume. It recovers the seam where every tested alternative fails, which we treat as evidence for the diagnosis rather than as a general-purpose method. On the BOSS-44 benchmark, the system improves full-chain success from 7.6% to 26.5%, a 3.5x improvement over the base policy and 51% of a privileged restoration oracle, whereas best-of-K resampling, a Diffusion Policy, and world-model baselines fail to recover from the evaluated seam states. On a real Franka arm running a fine-tuned π0.5π_{0.5} policy, the same monitor is limited by exterior-camera observability, yet closing the loop still recovers some otherwise-terminal failures, motivating wrist and gripper sensing. These results suggest that some long-horizon composition failures are better addressed by restoring the scene before resuming the policy than by retrying from an off-support state.

Figures & tables

Explore similar work

CardsList
  1. Behavior-Skill: A Fine-Grained Benchmark for Evaluating Vision-Language-Action Policies in Long-Horizon Tasks

    Aug 31, 2026Chunyun Ma, Lun Luo, Xingjian Luo +7Long-Horizon Robotic ManipulationMobile Manipulation

  2. ManiUnit: A Manipulation Skill Dataset and Benchmark for Long-Horizon Tasks

    Oct 8, 2026Guoting Wei, Dawei Yan, Xia Yuan +10Long-Horizon Robotic ManipulationRobot Manipulation Benchmarks