RoboAug: One Annotation to Hundreds of Scenes via Region-Contrastive Data Augmentation for Robotic Manipulation
Authors: Xinhua Wang, Kun Wu, Zhen Zhao, Hu Cao, Yinuo Zhao, Zhiyuan Xu, Meng Li, Shichao Fan, +5 more
Organizations: Beijing Innovation Center of Humanoid Robotics · Computation, Information and Technology, Technical University of Munich · City University of Hong Kong · The School of Mechanical Engineering and Automation, Beihang University · State Key Laboratory of Multimedia Information Processing, School of Computer Science, Peking University · The School of Advanced Manufacturing and Robotics, Peking University
Enhancing the generalization of robotic learning in diverse unseen environments remains a fundamental challenge. Existing approaches often rely on large-scale pretraining, which is labor-intensive and time-consuming, or semantic data augmentation methods that assume flawless upstream object detection in real-world scenarios. In this work, we propose RoboAug, a novel generative data augmentation framework that reduces reliance on large-scale pretraining and perfect visual recognition by requiring only a single image with bounding box annotations for dataset construction. Leveraging this minimal supervision, RoboAug employs pretrained generative models for precise semantic augmentation and introduces a plug-and-play region-contrastive loss to guide attention toward task-relevant regions, thereby enhancing generalization and task success rates. Extensive real-world experiments on UR-5e, AgileX, and Tian Gong 2.0 demonstrate that RoboAug consistently outperforms state-of-the-art augmentation baselines under background, distractor, and lighting shifts. Our project is available at https://x-roboaug.github.io/.
Figures & tables
Figure 2: Overview of RoboAug. RoboAug contains three stages: (1) task-relevant region extraction, (2) semantic data augmentation, and (3) region-contrastive policy learning.
Robot
Task
Traj.
Frame
Obj.
BBox.
Single-Arm Franka
8
2442
20511
19
87426
Single-Arm UR
17
4217
40136
34
197882
Dual-Arm UR
3
669
11077
11
71464
Dual-Arm Agilex
5
248
2025
9
10063
Total
33
7576
73749
46
366835
Table 1: RoboAug-D Dataset Statistics.
Figure 3: Comparison of [email protected] across RoboAug-D Dataset. We present the results of 5 representative objects.
Figure 4: Left: Overview of the generalization evaluation settings; Right: Experimental setup.
Method
UR-PutCornPot
UR-MoveLemon
Average
π0
0.30
0.35
-
π0 w/ RoboAug
0.65
0.70
-
Method
TG2-CollectBall
TG2-HeatBread
Average
π0
0.25
0.30
0.28
π0 w/ RoboAug
0.60
0.65
0.65
Table 2: Quantitative Results on π0 .
Augmentation Method
UR-PutCornPot
UR-MoveLemon
UR-StackBowl
UR-StoreCarrot
UR-OpenDrawerCorn
Average
ACT w/o Aug [ 75 ]
0.06
0.06
0.08
0.10
0.15
0.09
RoboEngine-T [ 71 ]
0.12
0.16
0.12
0.20
0.32
0.18
RoboEngine-G [ 71 ]
0.16
0.24
0.12
0.32
0.40
0.25
GenAug [ 12 ]
0.22
0.28
0.16
0.40
0.48
0.31
RoboAug
0.38
0.46
0.28
0.56
0.68
0.47
AGX-PutCornPlate
AGX-UprightMug
AGX-StackBowl
AGX-OpenPotCorn
AGX-CloseDrawerCorn
Table 3: Comparative results under triple-factor variations: 3 unseen backgrounds, 4 lighting conditions, and 3 distractors.
Method
RCL
UR-PutCornPot
AGX-PutCornPlate
TG2-WeightApple
Average
ACT w/o Aug [ 75 ]
×
0.06
0.12
0.20
0.13
ACT w/o Aug [ 75 ]
✓
0.12
0.20
0.32
0.21
RoboEngine-T [ 71 ]
×
0.12
0.24
0.28
0.21
RoboEngine-T [ 71 ]
✓
0.14
0.28
0.30
0.24
RoboEngine-G [ 71 ]
×
0.16
0.30
0.38
0.28
RoboEngine-G [ 71 ]
✓
0.16
0.36
0.44
0.32
Table 4: Ablation study results on region-contrastive loss (RCL).
Figure 5: Feature heatmap comparison of RoboAug with and without RCL.
Hyperparameter
Value
Hyperparameter
Value
ACT
Batch Size
24
π0
Batch Size
256
Learning Rate
1e-4
Learning Rate
5e-5
Optimizer
AdamW
Optimizer
AdamW
Vision Encoder
ResNet50
Vision Encoder
SigLip
Training Objective
ℓ1 + KL + RCL
Training Objective
Flow Matching + RCL
Training Step
50K
Training Step
30K
Table 5: Implementation Details.
Augmentation Method
UR-PutCornPlot
UR-MoveLemon
UR-StackBowl
UR-StoreCarrot
UR-OpenDrawerCorn
Average
No Aug
0.20
0.26
0.12
0.28
0.36
0.24
RoboEngine-T
0.25
0.38
0.20
0.36
0.55
0.35
RoboEngine-G
0.28
0.42
0.22
0.40
0.62
0.39
GenAug
0.30
0.44
0.20
0.46
0.68
0.42
RoboAug
0.36
0.52
0.28
0.58
0.84
0.52
AGX-PutCornPlate
AGX-UprightMug
AGX-StackBowl
AGX-OpenPotCorn
AGX-CloseDrawerCorn
Table 6: Quantitative results under the Dual-Factor Variation setting. We report the average success rates (%) on both UR-5e and AgileX robots. The evaluation involves 5 unseen background textures combined with 10 task-irrelevant distractors, totaling 100 trials for each task.
Figure 6: Background generalization on task UR-PutCornPot across 170 unseen backgrounds.
Augmentation Method
UR-PutCornPlot
UR-MoveLemon
UR-StackBowl
UR-StoreCarrot
UR-OpenDrawerCorn
Average
No Aug
0.25
0.20
0.30
0.10
0.25
0.22
RoboEngine-T
0.50
0.25
0.50
0.20
0.35
0.36
RoboEngine-G
0.55
0.30
0.50
0.20
0.40
0.39
GenAug
0.60
0.30
0.55
0.20
0.45
0.42
RoboAug
0.90
0.45
0.60
0.40
0.50
0.57
AGX-PutCornPlate
AGX-UprightMug
AGX-StackBowl
AGX-OpenPotCorn
AGX-CloseDrawerCorn
Table 7: Performance Comparison of different methods under a set of 10 distinct distractors.
Figure 7: Left: Comparison of task success rates between RoboAug and the best baseline method under varying numbers of distractors; Right: Comparison of RoboAug and best baseline method on 20 lighting conditions, evaluated using the task success rates.
Method
UR-PutCornPlot
UR-MoveLemon
UR-StackBowl
UR-StoreCarrot
UR-OpenDrawerCorn
Average
ACT w/o Aug
0.24
0.15
0.12
0.22
0.34
0.21
RoboEngine-T
0.24
0.26
0.22
0.20
0.40
0.26
RoboEngine-G
0.46
0.28
0.32
0.50
0.42
0.40
GenAug
0.48
0.36
0.28
0.46
0.42
0.40
RoboAug
0.90
0.60
0.60
0.65
0.65
0.68
AGX-PutCornPlate
AGX-UprightMug
AGX-StackBowl
AGX-OpenPotCorn
AGX-CloseDrawerCorn
Table 8: Tasks Success Rates on Novel Backgrounds. We evaluated the success rate of augmentation methods across 10 unseen backgrounds.
Figure 8: Effectiveness of RCL.
Figure 9: Experimental setup for evaluating generalization on the LIBERO-Plus benchmark.
Method
Background
Distractor
Light
Average
ACT w/o Aug [ 75 ]
0.745
0.860
0.789
0.798
RoboEngine-G [ 71 ]
0.806
0.942
0.855
0.868
RoboAug
0.913
0.990
0.896
0.933
Table 9: Generalization performance on the LIBERO-Plus benchmark.
Gen. Bg.
RCL
Ann.
SR ↑
No
No
0
0.15
No
Yes
1
0.33
Yes
No
1
0.55
Yes
Yes
1
0.68
Table 10: Component and mask-source ablations. (a) Contributions of background generation and RCL, using our masks when required. (b) Comparison of mask sources with both components enabled. Ann. denotes annotated frames per task; SR denotes average success rate. Oracle masks use annotations for all frames.
Reference condition
Success rate ↑
Clear
0.68
Partial occlusion
0.60
Unusual lighting
0.65
Rare object pose
0.68
Table 11: Sensitivity to reference-frame selection. Average success rates using reference frames with different visual conditions, with all other settings held fixed.
Attention variant
Success rate ↑
Full-image features (ours)
0.68
Masked-region features
0.15
Without attention gating
0.15
Table 12: Spatial attention ablation. Average success rates with attention computed from full-image features, attention computed from masked-region features, and no attention gating. All other settings are held fixed.
Figure 10: Effectiveness of data augmentation ratio.
Method
AGX-StackBowl
TG2-CollectBall
TG2-WeighApple
Average
Single-view w/o RoboAug
0.14
0.16
0.20
0.18
Single-view w RoboAug
0.62
0.52
0.80
0.65
Three-view w/o RoboAug
0.25
0.30
0.28
0.28
Three-view w RoboAug
0.70
0.64
0.85
0.73
Table 13: Multi-view augmentation performance across three tasks.
Figure 11: Visualization of the Regular Geometric background category.
Figure 12: Visualization of the Structured Design background category.
Figure 13: Visualization of the Scattered Pattern background category.
Figure 14: Visualization of the Scattered Pattern background category.
Figure 15: Visualization of the Scattered Pattern background category.
Figure 16: Visualization of the ten distinct distractors used in the UR and AgileX tasks.
Figure 17: Visualization of twenty distinct illumination conditions. The bottom row demonstrates dynamic lighting scenarios with multicolor changes at varying speeds.
The recent trend in scaling models for robot learning has resulted in impressive policies that can perform various manipulation tasks and generalize to novel scenarios. However, these policies continue to struggle with following instructions, likely due to the limited linguistic and action sequence diversity in existing robotics datasets. This paper introduces Task Robustness via Re-Labelling Vision-Action Robot Data (TREAD), a scalable framework that leverages large Vision-Language Models (VLMs) to augment existing robotics datasets without additional data collection, harnessing the transferable knowledge embedded in these models. Our approach leverages a pretrained VLM through three stages: generating semantic sub-tasks from original instruction labels and initial scenes, segmenting demonstration videos conditioned on these sub-tasks, and producing diverse instructions that incorporate object properties, effectively decomposing longer demonstrations into grounded language-action pairs. We further enhance robustness by augmenting the data with linguistically diverse versions of the text goals. Evaluations on LIBERO demonstrate that policies trained on our augmented datasets exhibit improved performance on novel, unseen tasks and goals. Our results show that TREAD enhances both planning generalization through trajectory decomposition and language-conditioned policy generalization through increased linguistic diversity.
Artur Kuramshin, Özgür Aslan, Cyrus Neary +1
1Mila — Quebec AI Institute · 2Université de Montréal · 3The University of British Columbia
Policies for robotic manipulation are produced by training on large teleoperated datasets. These datasets typically consist of free-space trajectories, making them difficult to transfer to test-time environments with obstacles. Previous methods for closing this gap have largely fallen into two groups. Dataset augmentation addresses it at training time but needs obstacle geometry in advance, whereas steering an existing checkpoint at inference time avoids that requirement but is limited in flexibility. Our method draws from both areas without inheriting either drawback. DetAug applies an obstacle-blind augmentation scheme to the transit phases of a free-space dataset, leaving object interactions untouched, and records the augmentation parameters as an explicit conditioning label. At inference it samples a batch of labels and executes the trajectory with the lowest collision cost. On the SafeLIBERO benchmark DetAug achieves a collision-free success rate more than 20pp above the next best method, and selecting over the label space outperforms guidance on the same policy by 26pp. On real hardware, inference-time steering methods collapse on tasks requiring large detours, while DetAug matches or exceeds an obstacle-conditioned baseline without ever seeing obstacles in training.
Reece O'Mahoney, Moritz Zoellner, Ioannis Havoutis
Oxford Robotics Institute, University of Oxford, Oxford, UK · Purdue University, West Lafayette, IN, USA
The scalability of robotic manipulation is fundamentally bottlenecked by the scarcity of task-aligned physical interaction data. While vision-language models (VLMs) and video generation models (VGMs) hold promise for autonomous data synthesis, they suffer from semantic-spatial misalignment and physical hallucinations, respectively. To bridge this gap, we introduce RoboEvolve, a novel framework that couples a VLM planner and a VGM simulator into a mutually reinforcing co-evolutionary loop. Operating purely on unlabeled seed images, RoboEvolve leverages a cognitive-inspired dual-phase mechanism: (i) daytime exploration fosters physically grounded behavioral discovery through a semantic-controlled multi-granular reward, and (ii) nighttime consolidation mines "near-miss" failures to stabilize policy optimization. Guided by an autonomous progressive curriculum, the system naturally scales from simple atomic actions to complex tasks. Extensive experiments demonstrate that RoboEvolve (I) achieves superior effectiveness, elevating base planners by 30 absolute points and amplifying simulator success by 48% on average; (II) exhibits extreme data efficiency, surpassing fully supervised baselines with merely 500 unlabeled seeds--a 50x reduction; and (III) demonstrates robust continual learning without catastrophic forgetting.
Harold Haodong Chen, Sirui Chen, Yingjie Xu +2
1The Hong Kong University of Science and Technology (Guangzhou) · 2The Hong Kong University of Science and Technology