Existing data generation methods for robot learning suffer from limited exploration, embodiment gaps, low signal-to-noise ratios, and domain shifts, leading to performance degradation during self-iteration and poor generalization to unseen scenes. To address these challenges, we propose Seed2Scale, a self-evolving data engine with parallel worlds expansion. Starting with as few as four seed demonstrations, Seed2Scale first executes a self-evolution stage driven by a heterogeneous synergy of "small-model collection, large-model evaluation, and target-model learning". Specifically, the lightweight Vision-Language-Action (VLA) model, SuperTiny, serves as a dedicated data collector for robust exploration. Concurrently, a pretrained Vision-Language Model (VLM) functions as a verifier to autonomously score and filter trajectories, supporting stable self-evolution in the evaluated tasks without performance collapse. Furthermore, Seed2Scale introduces a parallel worlds stage, projecting self-evolved trajectories into different environments of the same task to generate more diverse data and enhance adaptability to unseen scenes, including real-world environments. Experimental results demonstrate that Seed2Scale exhibits significant scaling potential: as iterations progress, the success rate of the target model shows a consistent upward trend, significantly outperforming the seed baseline. Notably, Seed2Scale achieves a remarkable 75.38% success rate in zero-shot real-world evaluations, where baseline methods fail completely (0%). Project page: https://terminators2025.github.io/Seed2Scale.github.io
Figures & tables
Fig. 1 : Architecture of the SuperTiny VLA. The model integrates vision (ResNet), language (T5), and robot state encodings into a unified conditioning memory Mt , which is processed by a lightweight Transformer decoder via cross-attention to predict action chunks. Temporal ensembling (Eq. 3 ) ensures smooth and consistent control.
Fig. 2 : (a) GR-1 robot tasks used for comparison with MimicGen and Seed2Scale ∗ . (b) Agibot A2 robot tasks for Seed2Scale evaluation.
Evaluated on Seen Scenes
Evaluated on Unseen Scenes
Task
SeedOnly
Seed2Scale ∗
Seed2Scale
SeedOnly
Seed2Scale ∗
Seed2Scale
Kitchen Cleanup
24.63%
71.43%
68.50%
0.00%
0.00%
75.51%
Cup-to-Cup Transfer
23.50%
64.14%
60.61%
0.00%
0.00%
61.62%
Can Stacking
7.50%
65.90%
67.70%
0.00%
0.00%
69.19%
Air Fryer Manipulation
33.08%
72.82%
71.21%
0.00%
0.00%
55.05%
Average
22.18%
68.57%
67.01%
0.00%
0.00%
65.34%
TABLE I : Success rate (%) of the target model on the four Agibot A2 tasks, evaluated on Seen and Unseen scenes over 396 independent trials, with best results per task in bold . For each variant, a single SmolVLA is trained on data combined from all four tasks and evaluated separately on each task.
Fig. 3 : (a) The original world, the parallel worlds projected from it, and the unseen worlds. (b) The worlds in the zero-shot real-world experiment.
Configuration
Episodes
Frames
Success Rate
SeedOnly
4
720
0.00%
Seed2Scale ∗
142
19,748
0.00%
Seed2Scale (3 scenes)
432
60,039
27.69%
Seed2Scale (5 scenes)
720
100,065
73.85%
Seed2Scale (7 scenes)
1,716
225,217
75.38%
TABLE II : Zero-shot success rate on the Pick-and-Place task, evaluated over 65 independent trials in an unseen real-world scene. Seed2Scale ( n scenes) denotes a configuration trained with parallel-world data from n of the 7 training scenes.
Fig. 4 : Scaling performance of the target model (SmolVLA) across self-evolution iterations.
Metric
Task
MimicGen
Seed2Scale ∗
Policy Succ.
Cylinder Grasp
37.25%
66.00%
Wheel Manip.
34.75%
93.25%
Average
36.00%
79.63%
Replay Succ.
Cylinder Grasp
21.00%
86.96%
Wheel Manip.
48.50%
67.86%
Average
34.75%
77.41%
TABLE III : Comparison with MimicGen on GR-1 tasks. We report the policy and replay success rates. Seed2Scale ∗ outperforms MimicGen on both.
Metric
Expert Demonstration
MimicGen
Seed2Scale ∗
Total Variation
1.32
3.68
1.34
Mean Absolute Jerk
0.0063
0.0261
0.0047
HF Power Ratio (%)
0.22
2.07
0.30
TABLE IV : Quantitative comparison of trajectory quality. Lower values indicate better performance. Seed2Scale ∗ produces trajectories with human-like smoothness, significantly outperforming MimicGen.
Fig. 5 : Comparison of robot action curves generated by expert demonstration, MimicGen, and the Seed2Scale ∗ .
Model
Params
Inference Time
Frequency
ACT
52M
45.67ms
21.9 Hz
Diffusion Policy
265M
135.83ms
7.4 Hz
SuperTiny
48M
38.08ms
26.3 Hz
TABLE V : Efficiency comparison of VLA models.
Fig. 6 : Performance comparison of SuperTiny, ACT, and Diffusion Policy as data collectors across self-evolution iterations. SuperTiny - denotes the variant without VLV-based quality filtering.
Fig. 7 : VLV analysis of successful trajectories with different quality levels. Top: regular-quality execution (moderate score). Bottom: high-quality execution (highest score).