SimVLA: Zero-Shot Sim-to-Real VLA Learning for Mobile Manipulation
Organizations: Yonsei University · Seoul National University
Abstract
Large-scale, diverse datasets have driven the success of LLMs and VLMs. But VLAs for robotics remain limited by the cost and complexity of real-world data collection. While simulation offers a scalable alternative, its potential for sim-to-real VLA learning in mobile manipulation remains largely underexplored. We introduce SimVLA, an end-to-end framework that trains VLAs entirely on synthetic simulation data without teleoperation for mobile manipulation. SimVLA is first pre-trained on two complementary simulation-derived datasets: SimAction, a large-scale robot action dataset spanning 35 diverse mobile manipulation tasks, generated by composing atomic skills, and SimVQA, which leverages privileged simulator state to provide spatial, geometric, and subtask-level visual-language supervision. We further post-train SimVLA on a mixture of SimAction and SimDeploy, a dataset collected from policy rollouts across diverse simulated environments. We evaluate SimVLA on tasks including restocking, pouring, and cleaning, and show zero-shot transfer to real-world mobile manipulation, including real home environments. SimVLA outperforms policies trained on 50 in-domain real-world demonstrations, suggesting that simulation can enable scalable sim-to-real mobile manipulation. We further demonstrate the value of multiple complementary forms of supervision for effectively leveraging simulation in VLA training.
Figures & tables
| Dataset | # Task | Mobility | V-L sup. | Data generation | Zero-shot Sim-to-Real |
|---|---|---|---|---|---|
| MimicGen ( Mandlekar et al., 2023 ) | 18 | ✗ | Teleoperation & Augmentation | ✗ | |
| RoboGen ( Wang et al., 2024 ) | 106 | ✗ | Motion planning & Optimization & RL | ✗ | |
| SkillMimicGen ( Garrett et al., 2025 ) | 18 | ✗ | ✗ | Teleoperation & Augmentation | ✗ |
| DexMimicGen ( Jiang et al., 2025 ) | 9 | ✗ | Teleoperation & Augmentation | ||
| GraspVLA ( Deng et al., 2025 ) | 1 | ✗ | Motion planning | ||
| RoboTwin 2.0 ( Chen et al., 2026 ) | 50 | ✗ | ✗ | Motion planning |
| MolmoB0T | SimVLA | |
| Pick | ||
| MolmoSpaces | 36.3 2.8 | 71.0 2.9 |
| SimVLA kitchens | 9.7 1.7 | 91.2 2.8 |
| PnP | ||
| MolmoSpaces | 1.0 0.7 | 49.0 3.4 |
| SimVLA kitchens | 0.5 0.5 | 79.5 3.1 |
Appendix figures & tables21 assets
Supplementary material from the paper’s appendix.
Appendix
| Objective | L2 mean (mm) | L2 max (mm) |
|---|---|---|
| None (URDF defaults) | ||
| Joint-space | ||
| EEF-space | ||
| Hybrid (joint + EEF) |
| ID | Task | Description |
|---|---|---|
| 1 | Move to mug | The robot navigates and positions itself near the mug. |
| 2 | Move to bowl | The robot navigates and positions itself near the bowl. |
| 3 | Move to bottle | The robot navigates and positions itself near the bottle. |
| 4 | Move to drawer | The robot navigates and positions itself in front of the drawer. |
| 5 | Put mug in sink | The robot navigates to the mug, grasps it, moves to the sink, and places it into the sink. |
| 6 | Put bowl in sink | The robot navigates to the bowl, grasps it, moves to the sink, and places it into the sink. |
| Domain | Randomness |
|---|---|
| Camera: Position Noise (m) | |
| Camera: Rotation Noise (deg) | |
| Kitchen: Types | 5 |
| Kitchen: Tables | 50 |
| Kitchen: Chairs | 31 |
| Kitchen: Cabinet Materials | 21 |
| VQA Function | Question Prompt Prototype | Answer Prototype |
|---|---|---|
| high_level_subtask | To complete the task “{task}”, what should the robot do now? | {subtask} |
| next_subtask | After current subtask {subtask}, what is the next subtask? | {subtask} |
| object_detection | Is the {object} visible in the front camera? | yes. bbox: {y1,x1,y2,x2} / no |
| object_3d_information | How far is the {object} from the robot base? | dx {x} dy {y} dz {z} |
| view_correspondence | Given front bbox {bbox}, what are the wrist camera bboxes? | left {bbox} right {bbox} |
| object_reachable | Is the robot close enough to manipulate the {object}? | yes / no. move base dx {x} dy {y} dyaw {yaw} |
| Hyperparameter | Pretrain | Post-train |
|---|---|---|
| Global batch size | 256 | 256 |
| Precision | BF16 | BF16 |
| Gradient clipping (L2) | 1.0 | 1.0 |
| Peak learning rate | ||
| End learning rate | ||
| LR scheduler | cosine | cosine |
| Hyperparameter | Value |
|---|---|
| FAST loss weight | 0.1 |
| VQA loss weight | 0.1 |
| FAST sampling ratio | 0.50 |
| Co-train sampling ratio | 0.01 |
| Discrete state input | false |
| Delta joint actions | true |
| Stage | Hardware | CPU | System RAM | Storage | Per-run compute | Total compute |
|---|---|---|---|---|---|---|
| Pre-training | 4 NVIDIA B200 (183GB) | 128 CPU cores | 1.5 TB | NVMe SSD | 12 days | 1152 GPU-hours |
| Post-training | 4 NVIDIA B200 (183GB) | 128 CPU cores | 1.5 TB | NVMe SSD | 30 hours | 120 GPU-hours |
| Model | Gripper | Subtask | 3D Pos. | Grasp | Cross-view IoU |
|---|---|---|---|---|---|
| L / R | Curr. / Next | (m) | (m) | ||
| SimVLA w/o SimVQA | – ∗ | – ∗ | – ∗ | – ∗ | – ∗ |
| SimVLA w/o TokSimAction | 0.86 / 0.69 | 0.18 / 0.39 | 0.50 | 0.16 | 0.04 |
| SimVLA w/o both | 0.85 / 0.68 | – † | – † | – † | – † |
| SimVLA | 1.00 / 0.89 | 0.69 / 0.77 | 0.06 | 0.05 | 0.61 |
| Task | Success Rate |
|---|---|
| Put mug to sink | 76% |
| Put bowl in drawer | 70% |
| Pour from bottle to mug | 76% |
| Put mug and bowl to sink | 38% |
| Put bowl to sink | 62% |
| Put bottle to sink | 64% |
| Model | Task 1 | Task 2 | Task 3 | Task 4 | Task 5 |
|---|---|---|---|---|---|
| fine-tuned w/ 50 demos | 0.55 0.43 [1pt] (8/20) | 0.36 0.39 [1pt] (5/20) | 0.48 0.38 [1pt] (6/20) | 0.45 0.46 [1pt] (7/20) | 0.43 0.47 [1pt] (7/20) |
| fine-tuned w/ 200 demos | 0.75 0.44 [1pt] (15/20) | 0.66 0.42 [1pt] (11/20) | 0.70 0.36 [1pt] (10/20) | 0.65 0.43 [1pt] (11/20) | 0.63 0.46 [1pt] (11/20) |
| SimVLA (w/o pre-train) | 0.28 0.38 [1pt] (3/20) | 0.15 0.16 [1pt] (0/20) | 0.15 0.25 [1pt] (1/20) | 0.18 0.34 [1pt] (2/20) | 0.18 0.29 [1pt] (1/20) |
| SimVLA (w/o SimDeploy) | 0.45 0.32 [1pt] (3/20) | 0.38 0.31 [1pt] (3/20) | 0.45 0.33 [1pt] (4/20) | 0.33 0.34 [1pt] (2/20) | 0.33 0.34 [1pt] (2/20) |
| SimVLA | 0.60 0.42 [1pt] (9/20) | 0.50 0.36 [1pt] (6/20) | 0.52 0.38 [1pt] (6/20) | 0.55 0.46 [1pt] (9/20) | 0.55 0.43 [1pt] (8/20) |
| Model | Task 1 | Task 2 | Task 3 | Task 4 | Task 5 |
|---|---|---|---|---|---|
| fine-tuned w/ 50 demos | 0.40 0.38 [1pt] (4/20) | 0.33 0.37 [1pt] (4/20) | 0.40 0.37 [1pt] (4/20) | 0.30 0.44 [1pt] (5/20) | 0.25 0.41 [1pt] (4/20) |
| fine-tuned w/ 200 demos | 0.45 0.39 [1pt] (5/20) | 0.43 0.40 [1pt] (6/20) | 0.48 0.35 [1pt] (4/20) | 0.35 0.46 [1pt] (6/20) | 0.28 0.44 [1pt] (5/20) |
| SimVLA (w/o pre-train) | 0.23 0.34 [1pt] (2/20) | 0.17 0.18 [1pt] (0/20) | 0.13 0.17 [1pt] (0/20) | 0.10 0.26 [1pt] (1/20) | 0.25 0.38 [1pt] (3/20) |
| SimVLA (w/o SimDeploy) | 0.45 0.36 [1pt] (4/20) | 0.36 0.31 [1pt] (3/20) | 0.43 0.38 [1pt] (5/20) | 0.38 0.32 [1pt] (2/20) | 0.50 0.32 [1pt] (4/20) |
| SimVLA | 0.58 0.41 [1pt] (8/20) | 0.50 0.36 [1pt] (6/20) | 0.52 0.37 [1pt] (6/20) | 0.60 0.45 [1pt] (10/20) | 0.68 0.41 [1pt] (11/20) |
| Model | Task 1 | Task 2 | Task 3 | Task 4 | Task 5 |
|---|---|---|---|---|---|
| fine-tuned w/ 50 demos | 0.35 0.46 [1pt] (6/20) | 0.37 0.39 [1pt] (5/20) | 0.40 0.37 [1pt] (4/20) | 0.50 0.49 [1pt] (9/20) | 0.30 0.41 [1pt] (4/20) |
| fine-tuned w/ 200 demos | 0.45 0.48 [1pt] (8/20) | 0.45 0.42 [1pt] (7/20) | 0.45 0.39 [1pt] (5/20) | 0.53 0.47 [1pt] (9/20) | 0.48 0.47 [1pt] (8/20) |
| SimVLA (w/o pre-train) | 0.20 0.34 [1pt] (2/20) | 0.13 0.15 [1pt] (0/20) | 0.13 0.17 [1pt] (0/20) | 0.18 0.29 [1pt] (1/20) | 0.08 0.18 [1pt] (0/20) |
| SimVLA (w/o SimDeploy) | 0.38 0.36 [1pt] (3/20) | 0.29 0.28 [1pt] (2/20) | 0.37 0.36 [1pt] (4/20) | 0.43 0.34 [1pt] (3/20) | 0.43 0.29 [1pt] (2/20) |
| SimVLA | 0.63 0.39 [1pt] (9/20) | 0.40 0.34 [1pt] (4/20) | 0.47 0.38 [1pt] (6/20) | 0.58 0.41 [1pt] (8/20) | 0.60 0.38 [1pt] (8/20) |
| Model | Task 1 | Task 2 | Task 3 | Task 4 | Task 5 |
|---|---|---|---|---|---|
| SimVLA w/o SimVQA | 0.43 0.47 [1pt] (7/20) | 0.36 0.41 [1pt] (5/20) | 0.37 0.40 [1pt] (5/20) | 0.50 0.43 [1pt] (7/20) | 0.45 0.43 [1pt] (6/20) |
| SimVLA w/o TokSimAction | 0.40 0.38 [1pt] (4/20) | 0.29 0.35 [1pt] (3/20) | 0.42 0.37 [1pt] (5/20) | 0.43 0.41 [1pt] (5/20) | 0.50 0.43 [1pt] (7/20) |
| SimVLA w/o both | 0.38 0.36 [1pt] (3/20) | 0.26 0.35 [1pt] (3/20) | 0.37 0.36 [1pt] (4/20) | 0.48 0.44 [1pt] (7/20) | 0.45 0.43 [1pt] (6/20) |
| SimVLA | 0.60 0.42 [1pt] (9/20) | 0.50 0.36 [1pt] (6/20) | 0.52 0.38 [1pt] (6/20) | 0.55 0.46 [1pt] (9/20) | 0.55 0.43 [1pt] (8/20) |
| Pantry (20 trials) | Real Home (10 trials) | ||||
| Model | Task 1 | Task 4 | Task 5 | Task 1 | Task 3 |
| fine-tuned w/ 50 Mock Kitchen demos | 0.10 0.21 [1pt] (0/20) | 0.08 0.18 [1pt] (0/20) | 0.18 0.24 [1pt] (0/20) | 0.20 0.26 [1pt] (0/10) | 0.17 0.18 [1pt] (0/10) |
| fine-tuned w/ 200 Mock Kitchen demos | 0.23 0.30 [1pt] (1/20) | 0.23 0.30 [1pt] (1/20) | 0.18 0.29 [1pt] (1/20) | 0.35 0.24 [1pt] (0/10) | 0.23 0.16 [1pt] (0/10) |
| SimVLA w/o pre-train | 0.10 0.21 [1pt] (0/20) | 0.10 0.26 [1pt] (1/20) | 0.08 0.18 [1pt] (0/20) | 0.10 0.21 [1pt] (0/10) | 0.10 0.16 [1pt] (0/10) |
| SimVLA | 0.60 0.38 [1pt] (8/20) | 0.53 0.47 [1pt] (9/20) | 0.58 0.44 [1pt] (9/20) | 0.70 0.42 [1pt] (6/10) | 0.47 0.42 [1pt] (3/10) |