cs.ROSep 29, 2026

FineART: Fine-Grained Annotated Robotic Trajectory Dataset and Vision-Language-Action Model for Bimanual Manipulation

Authors: Jade Choghari, Pepijn Kooijmans, Mansi Agarwal, Yusuf Umut Ciftci, Aseem Doriwala, Catherine Weaver, Mouli Sivapurapu, Kai Yang, +3 more

Organizations: Scale AI · Hugging Face · University of Southern California · Stanford University

Abstract

Robots operating in real-world environments must often execute complex, multi-step bimanual tasks over long horizons rather than single, isolated actions. Current manipulation datasets struggle to support this capability: although single-arm datasets reach hundreds of thousands of trajectories, they typically provide only one high-level instruction per episode, while existing bimanual datasets with subtask labels annotate only part of their recorded hours. We present FineART, a densely annotated bimanual manipulation dataset comprising 40,543 episodes (1,718 hours) and 533,913 subtasks across 151 tasks. We also introduce FineART-VLA, a vision-language-action policy that predicts its own next subtask to guide its actions. Mid-training on FineART's subtask annotations raises FineART-VLA's success at following spatial instructions from 32.0% to 100.0%. With step-by-step human subtask guidance, it also raises success on unseen long-horizon tasks from 16.0% to 76.0%. Furthermore, after minimal fine-tuning on a new robot, the policy requires only one-tenth of the data needed by baselines without this mid-training and generalizes zero-shot to tasks unseen on the new hardware. We open-source the full dataset, model weights, and training code.

Figures & tables

Appendix figures & tables9 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Scaling Bimanual Household Manipulation from 1,500 hours of Demonstrations to On-Policy Corrections

    Sep 3, 2026Jiafeng Xu, Qi Li, Yan Shen +7Bimanual ManipulationRobot Systems

  2. MonoDuo: Using One Robot Arm to Learn Bimanual Policies

    May 28, 2026Sandeep Bajamahal, Lawrence Yunliang Chen, Toru Lin +3Bimanual ManipulationRobotic Manipulation Policies

  3. Scalable Multi-Task Data Generation via Reinforcement Learning for Language-Conditioned Bimanual Dexterous Manipulation

    Jun 21, 2026Zechu Li, Yufeng Jin, Puze Liu +2Bimanual ManipulationDexterous Manipulation