Long-horizon robotic manipulation is often built by chaining independently trained skills. Although each skill can be reliable in isolation, performance degrades sharply when skills are chained: each downstream skill must start from the state its predecessor leaves behind rather than from its training distribution. We study this failure mode, Observation-Space Shift (OSS), and ask what causes these skill-seam failures. Using privileged simulator resets, we find that the dominant shift comes from displaced scene state (e.g., an open drawer or secondary objects left behind by earlier skills), not from the robot's joint configuration or the object the downstream skill manipulates. To test this diagnosis, we build a fully learned detect-restore-resume system: a task-progress monitor detects the stall, a learned policy restores the displaced scene components, and seam-robust fine-tuning lets the skill resume. It recovers the seam where every tested alternative fails, which we treat as evidence for the diagnosis rather than as a general-purpose method. On the BOSS-44 benchmark, the system improves full-chain success from 7.6% to 26.5%, a 3.5x improvement over the base policy and 51% of a privileged restoration oracle, whereas best-of-K resampling, a Diffusion Policy, and world-model baselines fail to recover from the evaluated seam states. On a real Franka arm running a fine-tuned π0.5 policy, the same monitor is limited by exterior-camera observability, yet closing the loop still recovers some otherwise-terminal failures, motivating wrist and gripper sensing. These results suggest that some long-horizon composition failures are better addressed by restoring the scene before resuming the policy than by retrying from an off-support state.
Figures & tables
Fig. 1: Detect, restore, resume at a skill seam. A chained skill collapses when it starts from the previous skill’s off-support terminal state; a task-progress monitor detects the stall, a learned restoration skill returns the displaced scene to support, and the chain resumes.
Sc
Skill
L0
Sc
Skill
L0
S1
open top drawer
1.00
S2
mid bowl on plate
0.45
S1
open bottom drawer
0.90
S2
mid bowl on cab. top
0.90
S1
bowl on plate
1.00
S4
close bottom drawer
1.00
S1
bowl on cab. top
0.70
S4
bowl in bottom drawer
0.80
S2
open top drawer
1.00
S4
bowl on cab. top
0.95
S2
bowl (back) on plate
0.60
S4
wine in bottom drawer
0.65
TABLE I: Per-skill success in isolation ( L0 ), catastrophic-chain skills. Mean 0.79 (range 0.35 – 1.00 ): the skills are competent alone, so chain collapse is a composition failure.
Fig. 2: The detect–restore–resume loop across scenes and on hardware. Each row runs one chained skill through a seam, left to right. Seam : the downstream skill begins from the previous skill’s off-support terminal state: an open drawer in simulation, a mid-place hand-off on hardware. Detect : the task-progress monitor fires (“Stalled!”). Restore : the learned restoration skill returns the displaced articulation and objects to the skill’s support. Resume : the skill resumes and completes ( ✓ ). Rows: two catastrophic simulation scenes (S1, open-drawer then place the bowl; S4, a distinct cabinet scene), then hardware chains RC1 (cup → plate) and RC3 (bowl → plate). Without the loop, execution halts at the Seam column. Simulation frames are the full system on the paper’s checkpoint (Table IV ); hardware shows recovered place attempts (Sec. V-F ). Each panel boxes the target object in red and shows a magnified inset of it. The box and inset turn green once the object is placed and the skill completes.
Restored component(s)
L2
Base (none)
0.098±0.031
Arm only
0.132±0.023
Arm + articulation
0.212±0.023
Arm + secondary objects
0.268±0.036
Full (oracle)
0.518±0.018
TABLE II: Component-wise reset ablation.
Alternative
Result at the OSS seam
Best-of- K ( K=8 )
0.00 (clean 0.35→1.00 )
Diffusion Policy ( 2500 ep)
0.00 (clean ≤0.70 )
World model (generator)
≈ base
World model (monitor)
AUROC <0.5
World model (planner)
0.65 vs. MB 0.53
End-to-end RL [ 21 ]
0.45 vs. 0.52
TABLE III: No tested alternative recovers the seam.
Method
L0
L1
L2
ΔL2 vs. Base ( p )
Base (no recovery)
0.885
0.343
0.076
n/a
Base ( + re-home/retry budget)
0.882
0.432
0.129
+0.053 (step effect)
Recovery only
0.876
0.459
0.175
+0.099 , p=0.094†
Prevention only
0.916
0.404
0.114
+0.038 , p=0.188
Prevention ( ∼ 6k seam demos)
0.898
0.353
0.117
+0.041 , p=0.062†
Recovery + Prev.
0.906
0.496
0.265
+0.189 , p=0.047∗
TABLE IV: Main result: four-condition ablation plus oracle ( 10 catastrophic chains, 5 seeds). L0 / L1 / L2 are success through one, two, and all three skills (the full chain); L2 is our primary metric, seed-averaged per chain. p is a sign-flip permutation test vs. Base over chains ( ∗p<0.05 ; †p<0.10 ), uncorrected across conditions, with the full-vs-base test pre-specified.
Regime
Chain
base
retry
restore
Restore decisive
ch3_2
0.00
0.00
0.27
ch3_3
0.01
0.12
0.29
ch3_11
0.00
0.06
0.09
Retry suffices
ch3_13
0.05
0.31
0.25
ch3_16
0.20
0.30
0.35
Restore hurts
ch3_8
0.32
0.30
0.21
TABLE V: When to restore vs. retry: per-chain L2 , recovery action isolated (matched re-home + retry vs. learned restoration, no prevention).
Fig. 3: The task-progress monitor in action (simulation, S1). Frames A–D from one execution with their monitor scores: reaching ( p=0.48 ), stalled at the seam ( p=0.70 ), reset to the seam by restoration ( p=0.00 ), and completed ( p=0.99 ). Below, the same scores over time: progress climbs, stalls ( Stall! ), the calibrated trigger fires (dashed), restoration drops it to zero, and the retry resumes. Monitor trigger + learned restoration, no oracle. Phases colored as in Fig. 1 .
Encoder / model
AUROC
Latency
Ours
DINOv3-B/16 (final)
0.903
6.1 ms
DINOv3-L/16
0.741
—
DINOv3-S/16
0.739
—
Wan-VAE (earlier)
0.782
38 ms
Baselines †
Robometer (4B PRM)
0.746
—
GVL (Qwen3-VL, 8B)
0.718
—
TABLE VI: Detector accuracy and cost (pooled seam-window AUROC).
Fig. 4: (a) The learned restoration skill reaches 1.00/0.93/0.78 (S1/S2/S4) at the skill level. (b) A distilled camera-only student matches the privileged teacher, removing simulator state from the loop.
Detection (exterior cam)
AUROC
95% CI
Leave-one-episode-out
0.625
[0.46,0.79]
Leave-one-chain-out
0.672
[0.51,0.83]
Recovery (matched flags)
Recovered
Baseline (no restore)
0/19
+ restoration
3/21
TABLE VII: Real-robot study: Franka + served π0.5 , three chains, 15 trials each.
Long-horizon household tasks require robots to compose many language-conditioned skills, yet the boundary between consecutive skills is rarely explicit. A skill may satisfy its own postcondition while leaving the robot, objects, or camera views in a state from which the next skill cannot reliably start. We study this semantic handoff problem in BEHAVIOR-1K through an agent-orchestrated vision-language-action execution harness. The harness invokes π0.5-based skill checkpoints trained from cleaned BEHAVIOR-1K demonstrations, assigns each skill typed arguments and a step budget, and uses multi-view vision-language model verification to decide whether execution should advance, retry, or replan. To separate isolated skill competence from long-horizon compositional robustness, we evaluate the same checkpoints under two initial-state distributions: clean skill-boundary snapshots and chained terminal states produced by previous skills. Selected navigation, grasping, placement, and door-opening skills achieve 77--100% success from clean snapshots under human-reviewed verification, yet composed rollouts still frequently stall from chained states. The resulting traces attribute failures to next-skill readiness, target grounding, and control execution, turning nearzero task success into actionable diagnostics for what VLA skill libraries must learn next: robustness to the messy chained-state distribution that clean demonstrations underrepresent.
Reliable execution of long-horizon mobile manipulation tasks remains challenging because overall task success depends on the successful completion of multiple constituent skills. Existing benchmarks, however, still rely primarily on full-task rollouts and aggregate task-level metrics, making intermediate failures difficult to observe and analyze. We present Behavior-Skill, a benchmark that reformulates the learning and evaluation of long-horizon tasks around executable constituent skills. It contains 235,492 skill instances from 10,000 demonstrations across 50 household tasks and 34 semantic skill categories. Each instance pairs a skill instruction with an aligned observation-action segment, and is further associated with a restorable intermediate state and a skill success condition to enable independent evaluation under valid preconditions. We further introduce trajectory-level and skill-level metrics to characterize policy capability beyond aggregate task success. Extensive experiments across representative VLA policies including pi0.5 and GR00T on the complete 50-task benchmark show that failures are highly non-uniform across skills, with contact-rich manipulation skills forming persistent bottlenecks. These results demonstrate that Behavior-Skill complements full-task evaluation by exposing intermediate capability profiles for analyzing and improving long-horizon VLA policies. Behavior-Skill is publicly available at https://github.com/nubot-nudt/Behavior-Skill.
Chunyun Ma, Lun Luo, Xingjian Luo +7
College of Intelligence Science and Technology and National Key Laboratory of Equipment State Sensing and Smart Support, National University of Defense Technology, Changsha, China. · XPeng Inc., Guangzhou, China. · The Chinese University of Hong Kong, Hong Kong, China. +1
Long-horizon mobile manipulation requires a robot to navigate multi-room environments and execute a sequence of manipulation skills under a single natural language instruction. Learning and evaluating these skills present three challenges: similar observations under a fixed task instruction may make skill selection ambiguous; even when a preceding skill succeeds, the robot state inherited by the next skill may deviate from its demonstrated starting states and affect execution; and task-level metrics hinder skill-specific diagnosis, while early failures leave later skills untested. We therefore introduce ManiUnit, a manipulation skill dataset and benchmark built from 50 BEHAVIOR-1K activities. Its dataset contains 137,899 segments across 21 skill types and 417 subtasks, and its benchmark contains 1,260 test instances. Correspondingly, ManiUnit pairs each segment with an explicit subtask instruction; measures sensitivity to perturbations of the robot's starting base position or joint configuration; and restores intermediate simulator states and defines local success conditions so that each skill can be evaluated without executing preceding stages. Evaluations of representative vision-language-action (VLA) policies show that similar aggregate scores can hide substantial per-skill differences. The tested starting-state perturbations also degrade execution: on the full benchmark, joint perturbations reduce success rates by approximately 56% relative to those from demonstrated starting states. On two long-horizon activities, a skill policy trained on ManiUnit segments achieves 78.7% local manipulation success, compared with 49.3% for a task policy trained on complete demonstrations. The trained skills further support complete-task execution on these activities, as coordinating the task and skill policies through a planner raises full-task success from 4.0% to 18.0%.
Guoting Wei, Dawei Yan, Xia Yuan +10
Nanjing University of Science and Technology · Intellifusion · Northwestern Polytechnical University