Existing data generation methods for robot learning suffer from limited exploration, embodiment gaps, low signal-to-noise ratios, and domain shifts, leading to performance degradation during self-iteration and poor generalization to unseen scenes. To address these challenges, we propose Seed2Scale, a self-evolving data engine with parallel worlds expansion. Starting with as few as four seed demonstrations, Seed2Scale first executes a self-evolution stage driven by a heterogeneous synergy of "small-model collection, large-model evaluation, and target-model learning". Specifically, the lightweight Vision-Language-Action (VLA) model, SuperTiny, serves as a dedicated data collector for robust exploration. Concurrently, a pretrained Vision-Language Model (VLM) functions as a verifier to autonomously score and filter trajectories, supporting stable self-evolution in the evaluated tasks without performance collapse. Furthermore, Seed2Scale introduces a parallel worlds stage, projecting self-evolved trajectories into different environments of the same task to generate more diverse data and enhance adaptability to unseen scenes, including real-world environments. Experimental results demonstrate that Seed2Scale exhibits significant scaling potential: as iterations progress, the success rate of the target model shows a consistent upward trend, significantly outperforming the seed baseline. Notably, Seed2Scale achieves a remarkable 75.38% success rate in zero-shot real-world evaluations, where baseline methods fail completely (0%). Project page: https://terminators2025.github.io/Seed2Scale.github.io
Figures & tables
Fig. 1 : Architecture of the SuperTiny VLA. The model integrates vision (ResNet), language (T5), and robot state encodings into a unified conditioning memory Mt , which is processed by a lightweight Transformer decoder via cross-attention to predict action chunks. Temporal ensembling (Eq. 3 ) ensures smooth and consistent control.
Fig. 2 : (a) GR-1 robot tasks used for comparison with MimicGen and Seed2Scale ∗ . (b) Agibot A2 robot tasks for Seed2Scale evaluation.
Evaluated on Seen Scenes
Evaluated on Unseen Scenes
Task
SeedOnly
Seed2Scale ∗
Seed2Scale
SeedOnly
Seed2Scale ∗
Seed2Scale
Kitchen Cleanup
24.63%
71.43%
68.50%
0.00%
0.00%
75.51%
Cup-to-Cup Transfer
23.50%
64.14%
60.61%
0.00%
0.00%
61.62%
Can Stacking
7.50%
65.90%
67.70%
0.00%
0.00%
69.19%
Air Fryer Manipulation
33.08%
72.82%
71.21%
0.00%
0.00%
55.05%
Average
22.18%
68.57%
67.01%
0.00%
0.00%
65.34%
TABLE I : Success rate (%) of the target model on the four Agibot A2 tasks, evaluated on Seen and Unseen scenes over 396 independent trials, with best results per task in bold . For each variant, a single SmolVLA is trained on data combined from all four tasks and evaluated separately on each task.
Fig. 3 : (a) The original world, the parallel worlds projected from it, and the unseen worlds. (b) The worlds in the zero-shot real-world experiment.
Configuration
Episodes
Frames
Success Rate
SeedOnly
4
720
0.00%
Seed2Scale ∗
142
19,748
0.00%
Seed2Scale (3 scenes)
432
60,039
27.69%
Seed2Scale (5 scenes)
720
100,065
73.85%
Seed2Scale (7 scenes)
1,716
225,217
75.38%
TABLE II : Zero-shot success rate on the Pick-and-Place task, evaluated over 65 independent trials in an unseen real-world scene. Seed2Scale ( n scenes) denotes a configuration trained with parallel-world data from n of the 7 training scenes.
Fig. 4 : Scaling performance of the target model (SmolVLA) across self-evolution iterations.
Metric
Task
MimicGen
Seed2Scale ∗
Policy Succ.
Cylinder Grasp
37.25%
66.00%
Wheel Manip.
34.75%
93.25%
Average
36.00%
79.63%
Replay Succ.
Cylinder Grasp
21.00%
86.96%
Wheel Manip.
48.50%
67.86%
Average
34.75%
77.41%
TABLE III : Comparison with MimicGen on GR-1 tasks. We report the policy and replay success rates. Seed2Scale ∗ outperforms MimicGen on both.
Metric
Expert Demonstration
MimicGen
Seed2Scale ∗
Total Variation
1.32
3.68
1.34
Mean Absolute Jerk
0.0063
0.0261
0.0047
HF Power Ratio (%)
0.22
2.07
0.30
TABLE IV : Quantitative comparison of trajectory quality. Lower values indicate better performance. Seed2Scale ∗ produces trajectories with human-like smoothness, significantly outperforming MimicGen.
Fig. 5 : Comparison of robot action curves generated by expert demonstration, MimicGen, and the Seed2Scale ∗ .
Model
Params
Inference Time
Frequency
ACT
52M
45.67ms
21.9 Hz
Diffusion Policy
265M
135.83ms
7.4 Hz
SuperTiny
48M
38.08ms
26.3 Hz
TABLE V : Efficiency comparison of VLA models.
Fig. 6 : Performance comparison of SuperTiny, ACT, and Diffusion Policy as data collectors across self-evolution iterations. SuperTiny - denotes the variant without VLV-based quality filtering.
Fig. 7 : VLV analysis of successful trajectories with different quality levels. Top: regular-quality execution (moderate score). Bottom: high-quality execution (highest score).
The scalability of robotic manipulation is fundamentally bottlenecked by the scarcity of task-aligned physical interaction data. While vision-language models (VLMs) and video generation models (VGMs) hold promise for autonomous data synthesis, they suffer from semantic-spatial misalignment and physical hallucinations, respectively. To bridge this gap, we introduce RoboEvolve, a novel framework that couples a VLM planner and a VGM simulator into a mutually reinforcing co-evolutionary loop. Operating purely on unlabeled seed images, RoboEvolve leverages a cognitive-inspired dual-phase mechanism: (i) daytime exploration fosters physically grounded behavioral discovery through a semantic-controlled multi-granular reward, and (ii) nighttime consolidation mines "near-miss" failures to stabilize policy optimization. Guided by an autonomous progressive curriculum, the system naturally scales from simple atomic actions to complex tasks. Extensive experiments demonstrate that RoboEvolve (I) achieves superior effectiveness, elevating base planners by 30 absolute points and amplifying simulator success by 48% on average; (II) exhibits extreme data efficiency, surpassing fully supervised baselines with merely 500 unlabeled seeds--a 50x reduction; and (III) demonstrates robust continual learning without catastrophic forgetting.
Harold Haodong Chen, Sirui Chen, Yingjie Xu +2
1The Hong Kong University of Science and Technology (Guangzhou) · 2The Hong Kong University of Science and Technology
Scaling robot learning requires large-scale, diverse demonstrations, yet real-world data collection via teleoperation remains prohibitively expensive and time-consuming. While video diffusion models offer a promising avenue for data scaling, existing generative approaches are often limited to superficial visual augmentation, or suffer from embodiment hallucinations that yield physically infeasible motions. We present a generalizable embodiment-centric world model that achieves scalable data generation by synthesizing photorealistic demonstrations with novel objects, in novel scenes, and from novel viewpoints. Our approach anchors generation to rendered robot motion while conditioning on explicit scene and object priors, effectively decoupling trajectory execution from environment synthesis. This formulation has the potential to unlock two powerful data scaling capabilities: (1) retrieval and rebirth, which repurposes existing trajectories into entirely new contexts without new motion data; and (2) prop-free teleoperation, where operators manipulate empty air and the model hallucinates the target objects and scene afterwards, eliminating reset time. We demonstrate with real-world experiments that our generated data consistently improves downstream policy performance and significantly reduces real-world data requirements across diverse manipulation tasks.
Junjie Ye, Rong Xue, Basile Van Hoorick +6
1USC Physical Superintelligence (PSI) Lab · 2Toyota Research Institute
The development of robust and generalizable robot learning models is critically contingent upon the availability of large-scale, diverse training data and reliable evaluation benchmarks. Collecting data in the physical world poses prohibitive costs and scalability challenges, and prevailing simulation benchmarks frequently suffer from fragmentation, narrow scope, or insufficient fidelity to enable effective sim-to-real transfer. To address these challenges, we introduce Genie Sim 3.0, a unified simulation platform for robotic manipulation. We present Genie Sim Generator, a large language model (LLM)-powered tool that constructs high-fidelity scenes from natural language instructions. Its principal strength resides in rapid and multi-dimensional generalization, facilitating the synthesis of diverse environments to support scalable data collection and robust policy evaluation. We introduce the first benchmark that pioneers the application of LLM for automated evaluation. It leverages LLM to mass-generate evaluation scenarios and employs Vision-Language Model (VLM) to establish an automated assessment pipeline. We also release an open-source dataset comprising more than 10,000 hours of synthetic data across over 200 tasks. Through systematic experimentation, we validate the robust zero-shot sim-to-real transfer capability of our open-source dataset, demonstrating that synthetic data can server as an effective substitute for real-world data under controlled conditions for scalable policy training. For code and dataset details, please refer to: https://github.com/AgibotTech/genie_sim.