Long-horizon robotic manipulation is often built by chaining independently trained skills. Although each skill can be reliable in isolation, performance degrades sharply when skills are chained: each downstream skill must start from the state its predecessor leaves behind rather than from its training distribution. We study this failure mode, Observation-Space Shift (OSS), and ask what causes these skill-seam failures. Using privileged simulator resets, we find that the dominant shift comes from displaced scene state (e.g., an open drawer or secondary objects left behind by earlier skills), not from the robot's joint configuration or the object the downstream skill manipulates. To test this diagnosis, we build a fully learned detect-restore-resume system: a task-progress monitor detects the stall, a learned policy restores the displaced scene components, and seam-robust fine-tuning lets the skill resume. It recovers the seam where every tested alternative fails, which we treat as evidence for the diagnosis rather than as a general-purpose method. On the BOSS-44 benchmark, the system improves full-chain success from 7.6% to 26.5%, a 3.5x improvement over the base policy and 51% of a privileged restoration oracle, whereas best-of-K resampling, a Diffusion Policy, and world-model baselines fail to recover from the evaluated seam states. On a real Franka arm running a fine-tuned π0.5 policy, the same monitor is limited by exterior-camera observability, yet closing the loop still recovers some otherwise-terminal failures, motivating wrist and gripper sensing. These results suggest that some long-horizon composition failures are better addressed by restoring the scene before resuming the policy than by retrying from an off-support state.
Figures & tables
Fig. 1: Detect, restore, resume at a skill seam. A chained skill collapses when it starts from the previous skill’s off-support terminal state; a task-progress monitor detects the stall, a learned restoration skill returns the displaced scene to support, and the chain resumes.
Sc
Skill
L0
Sc
Skill
L0
S1
open top drawer
1.00
S2
mid bowl on plate
0.45
S1
open bottom drawer
0.90
S2
mid bowl on cab. top
0.90
S1
bowl on plate
1.00
S4
close bottom drawer
1.00
S1
bowl on cab. top
0.70
S4
bowl in bottom drawer
0.80
S2
open top drawer
1.00
S4
bowl on cab. top
0.95
S2
bowl (back) on plate
0.60
S4
wine in bottom drawer
0.65
TABLE I: Per-skill success in isolation ( L0 ), catastrophic-chain skills. Mean 0.79 (range 0.35 – 1.00 ): the skills are competent alone, so chain collapse is a composition failure.
Fig. 2: The detect–restore–resume loop across scenes and on hardware. Each row runs one chained skill through a seam, left to right. Seam : the downstream skill begins from the previous skill’s off-support terminal state: an open drawer in simulation, a mid-place hand-off on hardware. Detect : the task-progress monitor fires (“Stalled!”). Restore : the learned restoration skill returns the displaced articulation and objects to the skill’s support. Resume : the skill resumes and completes ( ✓ ). Rows: two catastrophic simulation scenes (S1, open-drawer then place the bowl; S4, a distinct cabinet scene), then hardware chains RC1 (cup → plate) and RC3 (bowl → plate). Without the loop, execution halts at the Seam column. Simulation frames are the full system on the paper’s checkpoint (Table IV ); hardware shows recovered place attempts (Sec. V-F ). Each panel boxes the target object in red and shows a magnified inset of it. The box and inset turn green once the object is placed and the skill completes.
Restored component(s)
L2
Base (none)
0.098±0.031
Arm only
0.132±0.023
Arm + articulation
0.212±0.023
Arm + secondary objects
0.268±0.036
Full (oracle)
0.518±0.018
TABLE II: Component-wise reset ablation.
Alternative
Result at the OSS seam
Best-of- K ( K=8 )
0.00 (clean 0.35→1.00 )
Diffusion Policy ( 2500 ep)
0.00 (clean ≤0.70 )
World model (generator)
≈ base
World model (monitor)
AUROC <0.5
World model (planner)
0.65 vs. MB 0.53
End-to-end RL [ 21 ]
0.45 vs. 0.52
TABLE III: No tested alternative recovers the seam.
Method
L0
L1
L2
ΔL2 vs. Base ( p )
Base (no recovery)
0.885
0.343
0.076
n/a
Base ( + re-home/retry budget)
0.882
0.432
0.129
+0.053 (step effect)
Recovery only
0.876
0.459
0.175
+0.099 , p=0.094†
Prevention only
0.916
0.404
0.114
+0.038 , p=0.188
Prevention ( ∼ 6k seam demos)
0.898
0.353
0.117
+0.041 , p=0.062†
Recovery + Prev.
0.906
0.496
0.265
+0.189 , p=0.047∗
TABLE IV: Main result: four-condition ablation plus oracle ( 10 catastrophic chains, 5 seeds). L0 / L1 / L2 are success through one, two, and all three skills (the full chain); L2 is our primary metric, seed-averaged per chain. p is a sign-flip permutation test vs. Base over chains ( ∗p<0.05 ; †p<0.10 ), uncorrected across conditions, with the full-vs-base test pre-specified.
Regime
Chain
base
retry
restore
Restore decisive
ch3_2
0.00
0.00
0.27
ch3_3
0.01
0.12
0.29
ch3_11
0.00
0.06
0.09
Retry suffices
ch3_13
0.05
0.31
0.25
ch3_16
0.20
0.30
0.35
Restore hurts
ch3_8
0.32
0.30
0.21
TABLE V: When to restore vs. retry: per-chain L2 , recovery action isolated (matched re-home + retry vs. learned restoration, no prevention).
Fig. 3: The task-progress monitor in action (simulation, S1). Frames A–D from one execution with their monitor scores: reaching ( p=0.48 ), stalled at the seam ( p=0.70 ), reset to the seam by restoration ( p=0.00 ), and completed ( p=0.99 ). Below, the same scores over time: progress climbs, stalls ( Stall! ), the calibrated trigger fires (dashed), restoration drops it to zero, and the retry resumes. Monitor trigger + learned restoration, no oracle. Phases colored as in Fig. 1 .
Encoder / model
AUROC
Latency
Ours
DINOv3-B/16 (final)
0.903
6.1 ms
DINOv3-L/16
0.741
—
DINOv3-S/16
0.739
—
Wan-VAE (earlier)
0.782
38 ms
Baselines †
Robometer (4B PRM)
0.746
—
GVL (Qwen3-VL, 8B)
0.718
—
TABLE VI: Detector accuracy and cost (pooled seam-window AUROC).
Fig. 4: (a) The learned restoration skill reaches 1.00/0.93/0.78 (S1/S2/S4) at the skill level. (b) A distilled camera-only student matches the privileged teacher, removing simulator state from the loop.
Detection (exterior cam)
AUROC
95% CI
Leave-one-episode-out
0.625
[0.46,0.79]
Leave-one-chain-out
0.672
[0.51,0.83]
Recovery (matched flags)
Recovered
Baseline (no restore)
0/19
+ restoration
3/21
TABLE VII: Real-robot study: Franka + served π0.5 , three chains, 15 trials each.
College of Intelligence Science and Technology and National Key Laboratory of Equipment State Sensing and Smart Support, National University of Defense Technology, Changsha, China. · XPeng Inc., Guangzhou, China. · The Chinese University of Hong Kong, Hong Kong, China. +1