Robots operating in real-world environments must often execute complex, multi-step bimanual tasks over long horizons rather than single, isolated actions. Current manipulation datasets struggle to support this capability: although single-arm datasets reach hundreds of thousands of trajectories, they typically provide only one high-level instruction per episode, while existing bimanual datasets with subtask labels annotate only part of their recorded hours. We present FineART, a densely annotated bimanual manipulation dataset comprising 40,543 episodes (1,718 hours) and 533,913 subtasks across 151 tasks. We also introduce FineART-VLA, a vision-language-action policy that predicts its own next subtask to guide its actions. Mid-training on FineART's subtask annotations raises FineART-VLA's success at following spatial instructions from 32.0% to 100.0%. With step-by-step human subtask guidance, it also raises success on unseen long-horizon tasks from 16.0% to 76.0%. Furthermore, after minimal fine-tuning on a new robot, the policy requires only one-tenth of the data needed by baselines without this mid-training and generalizes zero-shot to tasks unseen on the new hardware. We open-source the full dataset, model weights, and training code.
Figures & tables
Figure 1 : Representative demonstrations from FineART dataset. FineART is the largest subtask-annotated manipulation dataset to date, consisting of 40.5K episodes and 534K subtasks, totaling 1,718 hours.
Dataset
Traj.
Tasks
Hours
Subtask Hours
Sec./Ep.
MT-Opt [ 19 ]
800,000
12
5,556
0
25
RH20T [ 12 ]
110,000
147
1,111
0
36.4
RoboSet [ 4 ]
7,500
38
N/A
0
N/A
BridgeData V2 [ 33 ]
60,096
13
127
0
7.6
Open X-Embodiment [ 28 ]
1M+
500+
N/A
0
N/A
DROID [ 20 ]
76,000
86
350
0
16.6
Table 1: Comparison of open-source robot manipulation datasets. FineART has the most densely annotated subtask hours. N/A: not reported.
Figure 2 : FineART dataset overview. (a) Distinct object classes represented in FineART and ABC-130K across semantic categories. (b) Frequency of descriptive object modifiers in subtask annotations. (c) Collected hours across manipulation primitives, highlighting the high-volume head and contact-rich tail. (d–f) Episode duration, subtask duration, and arm-speed distributions, respectively, compared with ABC-130K.
Figure 3 : FineART annotation schema. A multi-stage task with demonstration-level and temporally segmented subtask annotations. Subtask labels shown are illustrative categorical examples, not the actual annotations.
In-distribution
Partial ID
Out-of-distribution
Overall
Checkpoint
Shelf
Towel
Vase
Pegboard
Saucer
Sort tools
Avg.
(a) subtask training and knowledge insulation ( α=0.5 , full data)
Run 1.1 (subtask training + KI)
91.0 / 72.0
72.0 / 32.0
75.0 / 16.0
55.0 / 0.0
73.0 / 20.0
69.5 / 28.0
72.6 / 28.0
Run 1.2 (no subtask training + KI)
97.0 / 96.0
50.0 / 16.0
74.0 / 24.0
72.0 / 0.0
33.0 / 4.0
76.5 / 16.0
67.1 / 26.0
Run 1.3 (subtask training, no KI)
77.0 / 64.0
86.0 / 44.0
66.0 / 16.0
49.0 / 0.0
69.0 / 12.0
58.5 / 12.0
67.6 / 24.7
Run 1.4 (no subtask training, no KI)
82.0 / 72.0
58.5 / 4.0
41.0 / 12.0
65.0 / 0.0
58.0 / 12.0
50.0 / 4.0
59.1 / 17.3
Table 2 : In-embodiment (ALOHA) results on the 6-task evaluation suite.
Figure 4 : Success rate of FineART-VLA by task regime (ID, partial ID, OOD) with subtask supervision and knowledge insulation. ( n=25 per task, n=50 per regime, n=150 overall). Subtask training and KI improve most significantly on OOD tasks.
Figure 5 : Evaluation of subtask training on instruction following tasks (Run 1.3 against Run 1.4, no KI). The solid fill is success rate and the pale bar behind it progress rate.
Task
Behavior
Run 1.1 (subtask training + KI)
Sort tools
Corrected in-flight trajectory toward wrong bin
Run 1.2 (no subtask training + KI)
Sort tools
Recovery, and a handover between the arms
Sort tools
Handover between the arms
Run 1.4 (no subtask training, no KI)
Table 3: Notable behaviors observed in rollouts.
Figure 6 : FineART-VLA cross-embodiment YAM evaluation. Policies are mid-trained for 300k steps on the FineART, ABC, or MolmoAct2 datasets, then fine-tuned on YAM data for a fixed 5k steps at each data budget. Upstream π0.5 open-source weights are not fine-tuned in mid-trained checkpoints. Since ABC and MolmoAct2 datasets contain YAM data, zero-shot (ZS) policies are evaluated for ABC and MolmoAct2 only.
Number of Fine-Tuning Trajectories
Mid-Trained Checkpoint
Zero-shot
50
125
250
1250
(10/task)
(25/task)
(50/task)
(250/task)
(a) Five-task YAM benchmark
FineART-VLA (ours, ALOHA)
–
45.9 / 14.4
55.9 / 26.4
53.7 / 27.2
69.2 / 33.6
None (upstream π0.5 init)
–
25.0 / 2.4
33.4 / 6.4
36.1 / 8.0
50.8 / 18.4
ABC (XDOF, YAM data) ‡
22.5 / 0.0
63.0 / 24.8
66.0 / 25.6
73.1 / 34.4
64.7 / 24.0
Table 4: Cross-embodiment transfer to YAM, after a fixed 5 k-step fine-tune at each budget.
Figure 7 : FineART-VLA transfer to tasks absent from the YAM five-task fine-tuning set; 300k ALOHA mid-training steps + 5k fine-tuning steps on YAM.
Appendix figures & tables9 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 8 : FineART-VLA architecture. The backbone zθ is a 2 B-parameter Gemma language model with a SigLIP vision encoder, attending bidirectionally over the camera views and the task string. Its LM head decodes the subtask ℓsub and, right after it, the FAST-tokenized action sequence a~ . The separate 300 M-parameter action expert reads out the continuous chunk At through its own flow-matching loss. Under knowledge insulation, the action expert attends to stop-gradient copies of the backbone’s keys and values.
Setting
Value
(a) Mid-training data
Dataset
FineART, task-label-repaired release
Episodes / frames
40,543 / 185,534,500
Tasks
151 (dense task index)
Cameras
bird’s-eye, left wrist, right wrist ( 360×640 , resized 224×224 )
State / action
14 -dim bimanual joint + gripper, per frame
Appendix
Table 5 : Mid-training configuration for FineART-VLA.
Figure 9 : Block-causal attention pattern. Colored cells can attend to each other, white cells are masked, and the diagonal split means attention is causal within that block. (a) Generating the subtask ( 30% of samples). Images and the task string form the bidirectional prefix, and the subtask is generated one token at a time, so each subtask token can only see earlier subtask tokens – the autoregressive restriction described in Section 4.1 . (b) Predicting the action ( 70% of samples, or every sample for the no-subtask arm). Here the subtask and discretized state are given as input rather than generated, so they join the bidirectional prefix instead. The FAST-tokenized sequence is still generated token by token, so it stays causal like the subtask span in (a) . The continuous action tokens have no such restriction: flow matching denoises the whole chunk at once, so they can all attend freely to each other.
Figure 10 : Knowledge insulation at layer l . The VLM stream’s attention Attvlm(l) (top) is untouched by KI. The action stream’s attention Attact(l) (bottom) reads the VLM stream’s keys and values only through the stop-gradient copy sg(Kvlm,Vvlm) (purple, dashed), which blocks ∇θVLMLflow at the red × ( Eq. 6 ).
Figure 11 : Inference-time conditioning modes. The three panels show where the action expert’s low-level conditioning comes from. In (a) Flat it is just the raw task string, re-fed every chunk with the LM head never called. It is the only correct way to deploy the no-subtask arm (Run 1.4 in Table 2 ) since its LM head was never trained to produce anything. (b) Hierarchical is how the trained FineART-VLA is deployed ( Fig. 8 ): the LM head autoregressively generates a subtask ℓsub every N chunks and the action expert conditions on it until the next regeneration, which is how the subtask-trained arm (Run 1.3 ) is deployed. (c) Interactive skips the LM head altogether and lets an operator supply ℓsub directly. It is the human-oracle protocol behind the long-horizon instruction-following result in Fig. 5 , where the operator moves on to the next subtask once the previous one is judged done rather than waiting on the LM head to propose it.
Task
Total Rollouts
ALOHA Experiment
YAM Experiment
Put cup on shelf
575
Benchmark
Benchmark
Fold towel
575
Benchmark
Benchmark
Insert flower in vase
575
Benchmark
Benchmark
Insert tool into pegboard
575
Benchmark
Benchmark
Put cup on saucer
575
Benchmark
Benchmark
Sort tools
150
Benchmark
Cross-Embodiment Transfer
Appendix
Table 6 : Summary of Evaluations
Benchmark Suite
Benchmark Tasks
Initial Conditions
In-Embodiment
6
150
Cross-Embodiment
5
125
Total
11
275
Appendix
Table 7: Benchmark suites by initial conditions.
Figure 12 : Initial conditions for put cup on shelf.
Department of Computer Science, TU Darmstadt, Germany. · Honda Research Institute Europe GmbH, Offenbach, Germany. · DFKI, Research Department SAIROL, Darmstadt, Germany. +2