RoboAug: One Annotation to Hundreds of Scenes via Region-Contrastive Data Augmentation for Robotic Manipulation
Authors: Xinhua Wang, Kun Wu, Zhen Zhao, Hu Cao, Yinuo Zhao, Zhiyuan Xu, Meng Li, Shichao Fan, +5 more
Organizations: Beijing Innovation Center of Humanoid Robotics · Computation, Information and Technology, Technical University of Munich · City University of Hong Kong · The School of Mechanical Engineering and Automation, Beihang University · State Key Laboratory of Multimedia Information Processing, School of Computer Science, Peking University · The School of Advanced Manufacturing and Robotics, Peking University
Enhancing the generalization of robotic learning in diverse unseen environments remains a fundamental challenge. Existing approaches often rely on large-scale pretraining, which is labor-intensive and time-consuming, or semantic data augmentation methods that assume flawless upstream object detection in real-world scenarios. In this work, we propose RoboAug, a novel generative data augmentation framework that reduces reliance on large-scale pretraining and perfect visual recognition by requiring only a single image with bounding box annotations for dataset construction. Leveraging this minimal supervision, RoboAug employs pretrained generative models for precise semantic augmentation and introduces a plug-and-play region-contrastive loss to guide attention toward task-relevant regions, thereby enhancing generalization and task success rates. Extensive real-world experiments on UR-5e, AgileX, and Tian Gong 2.0 demonstrate that RoboAug consistently outperforms state-of-the-art augmentation baselines under background, distractor, and lighting shifts. Our project is available at https://x-roboaug.github.io/.
Figures & tables
Figure 2: Overview of RoboAug. RoboAug contains three stages: (1) task-relevant region extraction, (2) semantic data augmentation, and (3) region-contrastive policy learning.
Robot
Task
Traj.
Frame
Obj.
BBox.
Single-Arm Franka
8
2442
20511
19
87426
Single-Arm UR
17
4217
40136
34
197882
Dual-Arm UR
3
669
11077
11
71464
Dual-Arm Agilex
5
248
2025
9
10063
Total
33
7576
73749
46
366835
Table 1: RoboAug-D Dataset Statistics.
Figure 3: Comparison of [email protected] across RoboAug-D Dataset. We present the results of 5 representative objects.
Figure 4: Left: Overview of the generalization evaluation settings; Right: Experimental setup.
Method
UR-PutCornPot
UR-MoveLemon
Average
π0
0.30
0.35
-
π0 w/ RoboAug
0.65
0.70
-
Method
TG2-CollectBall
TG2-HeatBread
Average
π0
0.25
0.30
0.28
π0 w/ RoboAug
0.60
0.65
0.65
Table 2: Quantitative Results on π0 .
Augmentation Method
UR-PutCornPot
UR-MoveLemon
UR-StackBowl
UR-StoreCarrot
UR-OpenDrawerCorn
Average
ACT w/o Aug [ 75 ]
0.06
0.06
0.08
0.10
0.15
0.09
RoboEngine-T [ 71 ]
0.12
0.16
0.12
0.20
0.32
0.18
RoboEngine-G [ 71 ]
0.16
0.24
0.12
0.32
0.40
0.25
GenAug [ 12 ]
0.22
0.28
0.16
0.40
0.48
0.31
RoboAug
0.38
0.46
0.28
0.56
0.68
0.47
AGX-PutCornPlate
AGX-UprightMug
AGX-StackBowl
AGX-OpenPotCorn
AGX-CloseDrawerCorn
Table 3: Comparative results under triple-factor variations: 3 unseen backgrounds, 4 lighting conditions, and 3 distractors.
Method
RCL
UR-PutCornPot
AGX-PutCornPlate
TG2-WeightApple
Average
ACT w/o Aug [ 75 ]
×
0.06
0.12
0.20
0.13
ACT w/o Aug [ 75 ]
✓
0.12
0.20
0.32
0.21
RoboEngine-T [ 71 ]
×
0.12
0.24
0.28
0.21
RoboEngine-T [ 71 ]
✓
0.14
0.28
0.30
0.24
RoboEngine-G [ 71 ]
×
0.16
0.30
0.38
0.28
RoboEngine-G [ 71 ]
✓
0.16
0.36
0.44
0.32
Table 4: Ablation study results on region-contrastive loss (RCL).
Figure 5: Feature heatmap comparison of RoboAug with and without RCL.
Hyperparameter
Value
Hyperparameter
Value
ACT
Batch Size
24
π0
Batch Size
256
Learning Rate
1e-4
Learning Rate
5e-5
Optimizer
AdamW
Optimizer
AdamW
Vision Encoder
ResNet50
Vision Encoder
SigLip
Training Objective
ℓ1 + KL + RCL
Training Objective
Flow Matching + RCL
Training Step
50K
Training Step
30K
Table 5: Implementation Details.
Augmentation Method
UR-PutCornPlot
UR-MoveLemon
UR-StackBowl
UR-StoreCarrot
UR-OpenDrawerCorn
Average
No Aug
0.20
0.26
0.12
0.28
0.36
0.24
RoboEngine-T
0.25
0.38
0.20
0.36
0.55
0.35
RoboEngine-G
0.28
0.42
0.22
0.40
0.62
0.39
GenAug
0.30
0.44
0.20
0.46
0.68
0.42
RoboAug
0.36
0.52
0.28
0.58
0.84
0.52
AGX-PutCornPlate
AGX-UprightMug
AGX-StackBowl
AGX-OpenPotCorn
AGX-CloseDrawerCorn
Table 6: Quantitative results under the Dual-Factor Variation setting. We report the average success rates (%) on both UR-5e and AgileX robots. The evaluation involves 5 unseen background textures combined with 10 task-irrelevant distractors, totaling 100 trials for each task.
Figure 6: Background generalization on task UR-PutCornPot across 170 unseen backgrounds.
Augmentation Method
UR-PutCornPlot
UR-MoveLemon
UR-StackBowl
UR-StoreCarrot
UR-OpenDrawerCorn
Average
No Aug
0.25
0.20
0.30
0.10
0.25
0.22
RoboEngine-T
0.50
0.25
0.50
0.20
0.35
0.36
RoboEngine-G
0.55
0.30
0.50
0.20
0.40
0.39
GenAug
0.60
0.30
0.55
0.20
0.45
0.42
RoboAug
0.90
0.45
0.60
0.40
0.50
0.57
AGX-PutCornPlate
AGX-UprightMug
AGX-StackBowl
AGX-OpenPotCorn
AGX-CloseDrawerCorn
Table 7: Performance Comparison of different methods under a set of 10 distinct distractors.
Figure 7: Left: Comparison of task success rates between RoboAug and the best baseline method under varying numbers of distractors; Right: Comparison of RoboAug and best baseline method on 20 lighting conditions, evaluated using the task success rates.
Method
UR-PutCornPlot
UR-MoveLemon
UR-StackBowl
UR-StoreCarrot
UR-OpenDrawerCorn
Average
ACT w/o Aug
0.24
0.15
0.12
0.22
0.34
0.21
RoboEngine-T
0.24
0.26
0.22
0.20
0.40
0.26
RoboEngine-G
0.46
0.28
0.32
0.50
0.42
0.40
GenAug
0.48
0.36
0.28
0.46
0.42
0.40
RoboAug
0.90
0.60
0.60
0.65
0.65
0.68
AGX-PutCornPlate
AGX-UprightMug
AGX-StackBowl
AGX-OpenPotCorn
AGX-CloseDrawerCorn
Table 8: Tasks Success Rates on Novel Backgrounds. We evaluated the success rate of augmentation methods across 10 unseen backgrounds.
Figure 8: Effectiveness of RCL.
Figure 9: Experimental setup for evaluating generalization on the LIBERO-Plus benchmark.
Method
Background
Distractor
Light
Average
ACT w/o Aug [ 75 ]
0.745
0.860
0.789
0.798
RoboEngine-G [ 71 ]
0.806
0.942
0.855
0.868
RoboAug
0.913
0.990
0.896
0.933
Table 9: Generalization performance on the LIBERO-Plus benchmark.
Gen. Bg.
RCL
Ann.
SR ↑
No
No
0
0.15
No
Yes
1
0.33
Yes
No
1
0.55
Yes
Yes
1
0.68
Table 10: Component and mask-source ablations. (a) Contributions of background generation and RCL, using our masks when required. (b) Comparison of mask sources with both components enabled. Ann. denotes annotated frames per task; SR denotes average success rate. Oracle masks use annotations for all frames.
Reference condition
Success rate ↑
Clear
0.68
Partial occlusion
0.60
Unusual lighting
0.65
Rare object pose
0.68
Table 11: Sensitivity to reference-frame selection. Average success rates using reference frames with different visual conditions, with all other settings held fixed.
Attention variant
Success rate ↑
Full-image features (ours)
0.68
Masked-region features
0.15
Without attention gating
0.15
Table 12: Spatial attention ablation. Average success rates with attention computed from full-image features, attention computed from masked-region features, and no attention gating. All other settings are held fixed.
Figure 10: Effectiveness of data augmentation ratio.
Method
AGX-StackBowl
TG2-CollectBall
TG2-WeighApple
Average
Single-view w/o RoboAug
0.14
0.16
0.20
0.18
Single-view w RoboAug
0.62
0.52
0.80
0.65
Three-view w/o RoboAug
0.25
0.30
0.28
0.28
Three-view w RoboAug
0.70
0.64
0.85
0.73
Table 13: Multi-view augmentation performance across three tasks.
Figure 11: Visualization of the Regular Geometric background category.
Figure 12: Visualization of the Structured Design background category.
Figure 13: Visualization of the Scattered Pattern background category.
Figure 14: Visualization of the Scattered Pattern background category.
Figure 15: Visualization of the Scattered Pattern background category.
Figure 16: Visualization of the ten distinct distractors used in the UR and AgileX tasks.
Figure 17: Visualization of twenty distinct illumination conditions. The bottom row demonstrates dynamic lighting scenarios with multicolor changes at varying speeds.