ManiUnit: A Manipulation Skill Dataset and Benchmark for Long-Horizon Tasks
Organizations: Nanjing University of Science and Technology · Intellifusion · Northwestern Polytechnical University
Abstract
Long-horizon mobile manipulation requires a robot to navigate multi-room environments and execute a sequence of manipulation skills under a single natural language instruction. Learning and evaluating these skills present three challenges: similar observations under a fixed task instruction may make skill selection ambiguous; even when a preceding skill succeeds, the robot state inherited by the next skill may deviate from its demonstrated starting states and affect execution; and task-level metrics hinder skill-specific diagnosis, while early failures leave later skills untested. We therefore introduce ManiUnit, a manipulation skill dataset and benchmark built from 50 BEHAVIOR-1K activities. Its dataset contains 137,899 segments across 21 skill types and 417 subtasks, and its benchmark contains 1,260 test instances. Correspondingly, ManiUnit pairs each segment with an explicit subtask instruction; measures sensitivity to perturbations of the robot's starting base position or joint configuration; and restores intermediate simulator states and defines local success conditions so that each skill can be evaluated without executing preceding stages. Evaluations of representative vision-language-action (VLA) policies show that similar aggregate scores can hide substantial per-skill differences. The tested starting-state perturbations also degrade execution: on the full benchmark, joint perturbations reduce success rates by approximately 56% relative to those from demonstrated starting states. On two long-horizon activities, a skill policy trained on ManiUnit segments achieves 78.7% local manipulation success, compared with 49.3% for a task policy trained on complete demonstrations. The trained skills further support complete-task execution on these activities, as coordinating the task and skill policies through a planner raises full-task success from 4.0% to 18.0%.
Figures & tables
| Model | Original | Pos. | Color | Dist. |
|---|---|---|---|---|
| StarVLA-PI | 68.6% | 40.0% | 60.0% | 62.9% |
| StarVLA-GR00T | 71.4% | 20.0% | 65.7% | 62.9% |
| StarVLA-Cosmos2 | 60.0% | 22.9% | 54.3% | 60.0% |
| 91.4% | 74.3% | 80.0% | 80.0% |
| Task policy | Combined system | |||
|---|---|---|---|---|
| Activity | Full-task success | Mean Q-score | Full-task success | Mean Q-score |
| Radio | 4.0% | 4.0% | 24.0% | 24.0% |
| Trash | 4.0% | 18.0% | 12.0% | 26.0% |
| Overall | 4.0% | 11.0% | 18.0% | 25.0% |
| Dataset | Mobile | Training | Evaluation | ||
|---|---|---|---|---|---|
| Skill labels | Language instr. | Independent skill eval. | Robot-state perturb. | ||
| LIBERO ( Liu et al., 2023 ) | |||||
| CALVIN ( Mees et al., 2022 ) | |||||
| RoboTwin 2.0 ( Chen et al., 2026 ) | |||||
| RoboCasa365 ( Nasiriany et al., 2026 ) | |||||
| BEHAVIOR-1K ( Li et al., 2023 ) | |||||
Appendix figures & tables11 assets
Supplementary material from the paper’s appendix.
Appendix
| Skill type | Source activities | Subtasks | Segments |
|---|---|---|---|
| Pick Up | 50 | 153 | 52,251 |
| Place | 45 | 184 | 49,574 |
| Chop | 6 | 18 | 7,624 |
| Open Door | 24 | 9 | 7,476 |
| Close Door | 20 | 7 | 6,048 |
| Pour | 7 | 14 | 2,993 |
| Parameter | Value |
|---|---|
| Monitored arm joints | First 3 per arm |
| Reference window | 500 frames |
| Smoothing window | 5 frames |
| Minimum joint departure | |
| Minimum qualifying duration | 5 frames |
| Gripper position-change threshold |
| StarVLA-PI | StarVLA-GR00T | |||||||
|---|---|---|---|---|---|---|---|---|
| Skill | Original | Base | Joint | Mean | Original | Base | Joint | Mean |
| Attach | 21.7% | 11.7% | 3.3% | 12.2% | 26.7% | 1.7% | 0.0% | 9.4% |
| Chop | 41.7% | 31.7% | 46.7% | 40.0% | 43.3% | 31.7% | 41.7% | 38.9% |
| Close Door | 83.3% | 83.3% | 50.0% | 72.2% | 90.0% | 83.3% | 46.7% | 73.3% |
| Close Drawer | 86.7% | 95.0% | 35.0% | 72.2% | 98.3% | 98.3% | 46.7% | 81.1% |
| Close Lid | 73.3% | 66.7% | 31.7% | 57.2% | 75.0% | 61.7% | 35.0% | 57.2% |
| StarVLA-PI | StarVLA-GR00T | |||||
| Skill | Original | Base | Joint | Original | Base | Joint |
| Close Door | 95.0% | 97.5% | 45.0% | 95.0% | 90.0% | 42.5% |
| Open Door | 80.0% | 96.7% | 13.3% | 93.3% | 83.3% | 10.0% |
| Pick Up | 68.6% | 77.1% | 8.6% | 71.4% | 80.0% | 14.3% |
| Place | 91.4% | 74.3% | 80.0% | 85.7% | 57.1% | 68.6% |
| Pour | 22.9% | 5.7% | 8.6% | 57.1% | 25.7% | 37.1% |