Robots operating in real-world environments must often execute complex, multi-step bimanual tasks over long horizons rather than single, isolated actions. Current manipulation datasets struggle to support this capability: although single-arm datasets reach hundreds of thousands of trajectories, they typically provide only one high-level instruction per episode, while existing bimanual datasets with subtask labels annotate only part of their recorded hours. We present FineART, a densely annotated bimanual manipulation dataset comprising 40,543 episodes (1,718 hours) and 533,913 subtasks across 151 tasks. We also introduce FineART-VLA, a vision-language-action policy that predicts its own next subtask to guide its actions. Mid-training on FineART's subtask annotations raises FineART-VLA's success at following spatial instructions from 32.0% to 100.0%. With step-by-step human subtask guidance, it also raises success on unseen long-horizon tasks from 16.0% to 76.0%. Furthermore, after minimal fine-tuning on a new robot, the policy requires only one-tenth of the data needed by baselines without this mid-training and generalizes zero-shot to tasks unseen on the new hardware. We open-source the full dataset, model weights, and training code.
Figures & tables
Figure 1 : Representative demonstrations from FineART dataset. FineART is the largest subtask-annotated manipulation dataset to date, consisting of 40.5K episodes and 534K subtasks, totaling 1,718 hours.
Dataset
Traj.
Tasks
Hours
Subtask Hours
Sec./Ep.
MT-Opt [ 19 ]
800,000
12
5,556
0
25
RH20T [ 12 ]
110,000
147
1,111
0
36.4
RoboSet [ 4 ]
7,500
38
N/A
0
N/A
BridgeData V2 [ 33 ]
60,096
13
127
0
7.6
Open X-Embodiment [ 28 ]
1M+
500+
N/A
0
N/A
DROID [ 20 ]
76,000
86
350
0
16.6
Table 1: Comparison of open-source robot manipulation datasets. FineART has the most densely annotated subtask hours. N/A: not reported.
Figure 2 : FineART dataset overview. (a) Distinct object classes represented in FineART and ABC-130K across semantic categories. (b) Frequency of descriptive object modifiers in subtask annotations. (c) Collected hours across manipulation primitives, highlighting the high-volume head and contact-rich tail. (d–f) Episode duration, subtask duration, and arm-speed distributions, respectively, compared with ABC-130K.
Figure 3 : FineART annotation schema. A multi-stage task with demonstration-level and temporally segmented subtask annotations. Subtask labels shown are illustrative categorical examples, not the actual annotations.
In-distribution
Partial ID
Out-of-distribution
Overall
Checkpoint
Shelf
Towel
Vase
Pegboard
Saucer
Sort tools
Avg.
(a) subtask training and knowledge insulation ( α=0.5 , full data)
Run 1.1 (subtask training + KI)
91.0 / 72.0
72.0 / 32.0
75.0 / 16.0
55.0 / 0.0
73.0 / 20.0
69.5 / 28.0
72.6 / 28.0
Run 1.2 (no subtask training + KI)
97.0 / 96.0
50.0 / 16.0
74.0 / 24.0
72.0 / 0.0
33.0 / 4.0
76.5 / 16.0
67.1 / 26.0
Run 1.3 (subtask training, no KI)
77.0 / 64.0
86.0 / 44.0
66.0 / 16.0
49.0 / 0.0
69.0 / 12.0
58.5 / 12.0
67.6 / 24.7
Run 1.4 (no subtask training, no KI)
82.0 / 72.0
58.5 / 4.0
41.0 / 12.0
65.0 / 0.0
58.0 / 12.0
50.0 / 4.0
59.1 / 17.3
Table 2 : In-embodiment (ALOHA) results on the 6-task evaluation suite.
Figure 4 : Success rate of FineART-VLA by task regime (ID, partial ID, OOD) with subtask supervision and knowledge insulation. ( n=25 per task, n=50 per regime, n=150 overall). Subtask training and KI improve most significantly on OOD tasks.
Figure 5 : Evaluation of subtask training on instruction following tasks (Run 1.3 against Run 1.4, no KI). The solid fill is success rate and the pale bar behind it progress rate.
Task
Behavior
Run 1.1 (subtask training + KI)
Sort tools
Corrected in-flight trajectory toward wrong bin
Run 1.2 (no subtask training + KI)
Sort tools
Recovery, and a handover between the arms
Sort tools
Handover between the arms
Run 1.4 (no subtask training, no KI)
Table 3: Notable behaviors observed in rollouts.
Figure 6 : FineART-VLA cross-embodiment YAM evaluation. Policies are mid-trained for 300k steps on the FineART, ABC, or MolmoAct2 datasets, then fine-tuned on YAM data for a fixed 5k steps at each data budget. Upstream π0.5 open-source weights are not fine-tuned in mid-trained checkpoints. Since ABC and MolmoAct2 datasets contain YAM data, zero-shot (ZS) policies are evaluated for ABC and MolmoAct2 only.
Number of Fine-Tuning Trajectories
Mid-Trained Checkpoint
Zero-shot
50
125
250
1250
(10/task)
(25/task)
(50/task)
(250/task)
(a) Five-task YAM benchmark
FineART-VLA (ours, ALOHA)
–
45.9 / 14.4
55.9 / 26.4
53.7 / 27.2
69.2 / 33.6
None (upstream π0.5 init)
–
25.0 / 2.4
33.4 / 6.4
36.1 / 8.0
50.8 / 18.4
ABC (XDOF, YAM data) ‡
22.5 / 0.0
63.0 / 24.8
66.0 / 25.6
73.1 / 34.4
64.7 / 24.0
Table 4: Cross-embodiment transfer to YAM, after a fixed 5 k-step fine-tune at each budget.
Figure 7 : FineART-VLA transfer to tasks absent from the YAM five-task fine-tuning set; 300k ALOHA mid-training steps + 5k fine-tuning steps on YAM.
Appendix figures & tables9 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 8 : FineART-VLA architecture. The backbone zθ is a 2 B-parameter Gemma language model with a SigLIP vision encoder, attending bidirectionally over the camera views and the task string. Its LM head decodes the subtask ℓsub and, right after it, the FAST-tokenized action sequence a~ . The separate 300 M-parameter action expert reads out the continuous chunk At through its own flow-matching loss. Under knowledge insulation, the action expert attends to stop-gradient copies of the backbone’s keys and values.
Setting
Value
(a) Mid-training data
Dataset
FineART, task-label-repaired release
Episodes / frames
40,543 / 185,534,500
Tasks
151 (dense task index)
Cameras
bird’s-eye, left wrist, right wrist ( 360×640 , resized 224×224 )
State / action
14 -dim bimanual joint + gripper, per frame
Appendix
Table 5 : Mid-training configuration for FineART-VLA.
Figure 9 : Block-causal attention pattern. Colored cells can attend to each other, white cells are masked, and the diagonal split means attention is causal within that block. (a) Generating the subtask ( 30% of samples). Images and the task string form the bidirectional prefix, and the subtask is generated one token at a time, so each subtask token can only see earlier subtask tokens – the autoregressive restriction described in Section 4.1 . (b) Predicting the action ( 70% of samples, or every sample for the no-subtask arm). Here the subtask and discretized state are given as input rather than generated, so they join the bidirectional prefix instead. The FAST-tokenized sequence is still generated token by token, so it stays causal like the subtask span in (a) . The continuous action tokens have no such restriction: flow matching denoises the whole chunk at once, so they can all attend freely to each other.
Figure 10 : Knowledge insulation at layer l . The VLM stream’s attention Attvlm(l) (top) is untouched by KI. The action stream’s attention Attact(l) (bottom) reads the VLM stream’s keys and values only through the stop-gradient copy sg(Kvlm,Vvlm) (purple, dashed), which blocks ∇θVLMLflow at the red × ( Eq. 6 ).
Figure 11 : Inference-time conditioning modes. The three panels show where the action expert’s low-level conditioning comes from. In (a) Flat it is just the raw task string, re-fed every chunk with the LM head never called. It is the only correct way to deploy the no-subtask arm (Run 1.4 in Table 2 ) since its LM head was never trained to produce anything. (b) Hierarchical is how the trained FineART-VLA is deployed ( Fig. 8 ): the LM head autoregressively generates a subtask ℓsub every N chunks and the action expert conditions on it until the next regeneration, which is how the subtask-trained arm (Run 1.3 ) is deployed. (c) Interactive skips the LM head altogether and lets an operator supply ℓsub directly. It is the human-oracle protocol behind the long-horizon instruction-following result in Fig. 5 , where the operator moves on to the next subtask once the previous one is judged done rather than waiting on the LM head to propose it.
Task
Total Rollouts
ALOHA Experiment
YAM Experiment
Put cup on shelf
575
Benchmark
Benchmark
Fold towel
575
Benchmark
Benchmark
Insert flower in vase
575
Benchmark
Benchmark
Insert tool into pegboard
575
Benchmark
Benchmark
Put cup on saucer
575
Benchmark
Benchmark
Sort tools
150
Benchmark
Cross-Embodiment Transfer
Appendix
Table 6 : Summary of Evaluations
Benchmark Suite
Benchmark Tasks
Initial Conditions
In-Embodiment
6
150
Cross-Embodiment
5
125
Total
11
275
Appendix
Table 7: Benchmark suites by initial conditions.
Figure 12 : Initial conditions for put cup on shelf.
Learning generalist policies for robust bimanual manipulation is bottlenecked by the scarcity of high quality large scale human demonstration data. In this work, we release 1,500 hours of diverse bimanual manipulation demonstrations covering everyday household tasks, and use this comprehensive corpus to train XR-2, a powerful vision-language-action (VLA) model. Enabled by a purpose built high throughput data pipeline and a carefully designed multi stage training paradigm, XR-2 attains strong manipulation performance in our systematic experiments while retaining favorable training efficiency and high data utilization. We further study two critical scaling axes: varying the amount of expert demonstration data, and post training on DAgger correction data from real time human interventions. In both settings, task success rate improves steadily over the data ranges we probe, exhibiting a clear consistent scaling trend at our current data scale. These results validate both the learning capacity of XR-2 and the promising scaling properties of the released dataset, which we open source to support reproducible research on bimanual robot manipulation learning.
Jiafeng Xu, Qi Li, Yan Shen +7
School of Computer Science, Peking University. · 1PrimeBot Research Institute, Swancor Advanced Materials Co., Ltd. · 3Crobotia.
Bimanual coordination is essential for many real-world manipulation tasks, yet learning bimanual robot policies is limited by the scarcity of bimanual robots and datasets. Single-arm robots, however, are widely available in research labs. Can we leverage them to train bimanual robot policies? We present MonoDuo, a framework for learning bimanual manipulation policies using single-arm robot demonstrations paired with human collaboration. MonoDuo collects data by teleoperating a single-arm robot to perform one side of a bimanual task while a human performs the other, then swapping roles to cover both sides. RGB-D observations from a wrist-mounted and fixed camera are augmented into synthetic demonstrations for target bimanual robots using state-of-the-art hand pose estimation, image and point cloud segmentation, and inpainting. These synthetic demonstrations, grounded in real robot kinematics, are used to train bimanual policies. We evaluate MonoDuo on five tasks: box lifting, backpack packing, cloth folding, jacket zipping, and plate handover. Compared to approaches relying solely on human bimanual videos, MonoDuo enables zero-shot deployment on unseen bimanual robot configurations, achieving success rates up to 70%. With only 25 target robot demonstrations, few-shot finetuning further boosts success rates by 65-70% over training from scratch, demonstrating MonoDuo's effectiveness in efficiently transferring knowledge from single-arm robot data to bimanual robot policies.
Sandeep Bajamahal, Lawrence Yunliang Chen, Toru Lin +3
A key bottleneck in training generalist policies for bimanual dexterous manipulation is the lack of large-scale, high-quality datasets. Synthetic data generation in simulation provides a scalable alternative to human video demonstrations by overcoming challenges such as morphology mismatch, missing physical interactions, and the generation of robot actions. However, existing approaches based on human teleoperation offer limited task diversity, as object-centric trajectory matching often neglects the feasibility of robot execution. Reinforcement learning (RL) enables broader scalability but is often constrained by handcrafted, task-specific rewards. In this work, we propose a systematic RL-based data generation pipeline that integrates generalizable reward design, effective domain randomization, and language-conditioned task annotations. This pipeline synthesizes diverse, high-quality datasets for dexterous bimanual manipulation and enables training of language-conditioned multi-task policies. Our experiments show that the generated data significantly improves generalization across three representative manipulation tasks.
Zechu Li, Yufeng Jin, Puze Liu +2
Department of Computer Science, TU Darmstadt, Germany. · Honda Research Institute Europe GmbH, Offenbach, Germany. · DFKI, Research Department SAIROL, Darmstadt, Germany. +2