Scaling robot data is crucial for building generalist Vision-Language-Action (VLA) models, yet robot trajectories are harder to scale than web-scale image-text data because embodied collection is costly and sparsely covers the physical world. This makes representation quality a central bottleneck: under a fixed robot-data budget, continued pre-training must turn limited trajectories into transferable visual-action knowledge rather than merely fit actions. We propose VLAct, a VLA-oriented VLM backbone trained on broad, heterogeneous, multi-embodiment robot data before task-specific fine-tuning. VLAct preserves the broad VLM prior and encourages shared action semantics across embodiments through VLM-prior preservation, multi-head continuous action co-supervision, and a partially unified cross-embodiment action layout, while allowing task-specific action heads during fine-tuning. Across simulation, real-world, and unseen-embodiment transfer, VLAct consistently improves downstream performance under fixed fine-tuning protocols. On LIBERO-Plus and RoboTwin 2.0, VLAct surpasses industrial VLA systems including ABot-M0 and LingBot-VLA, achieving success rates of 82.6% and 92.5%. On RoboDojo, VLAct ranks sixth among all policies by success rate and outperforms all explicitly designated world-action model (WAM) entries on both metrics. Most notably, on RoboCasa-GR1, an unseen humanoid embodiment, VLAct using only 20% of downstream trajectories outperforms the full-data GR00T-N1.6 baseline. These results are obtained using fully open-source data and only a 16-GPU training setup, showing that representation-centric continued pre-training can deliver highly competitive performance under a modest compute budget and is an important independent axis of VLA progress beyond data scaling.
Figures & tables
Method
Camera
Robot
Lang.
Light
Bg.
Noise
Layout
Total
OpenVLA [ 54 ]
0.8
3.5
23.0
8.1
34.8
15.2
28.5
15.6
OpenVLA-OFT [ 53 ]
56.4
31.9
79.5
88.7
93.3
75.8
74.2
69.6
NORA [ 44 ]
2.2
37.0
65.1
45.7
58.6
12.8
62.1
39.0
WorldVLA [ 19 ]
0.1
27.9
41.6
43.7
17.1
10.9
38.0
25.0
UniVLA [ 90 ]
1.8
46.2
69.6
69.0
81.0
21.2
31.9
42.9
π0 [ 12 ]
13.8
6.0
58.8
85.0
81.4
79.0
68.8
53.6
Table 1 : Results on LIBERO-Plus. Per-dimension success rate (%) on the LIBERO-Plus robustness benchmark. All models are trained on standard LIBERO and evaluated on the perturbed test set. Dark blue highlights VLAct and per-column best results; light blue highlights the strongest prior method by total score and per-column second-best results.
Method
Family
Base Setting
Data Scaling Setting
Clean
Clean
Random
Published VLA and world-action-model systems
π0 [ 12 ]
Flow
46.4
65.9
58.4
π0.5 [ 45 , 95 ]
Flow
60.2
82.7
76.8
X-VLA [ 114 , 10 ]
Flow
70.0
72.8
72.8
Lingbot-VLA [ 95 ]
Flow
–
88.6
86.7
Table 2: RoboTwin 2.0 results. Task success rate (%) under the Base and Data Scaling settings. Base uses 50 clean trajectories per task and is evaluated on the Clean regime; Data Scaling additionally uses 500 randomized trajectories per task and is evaluated on both the Clean and Random regimes. Each reported regime evaluates 50 tasks with 100 episodes per task. Published numbers are taken from the cited papers or their same-protocol re-evaluations. Dark blue highlights our default VLAct-OFT configuration and light blue highlights alternative VLAct heads. Best results are bold and second-best results are underlined.
Method
Gen.-Std.
Gen.-Rand.
Precision
Long-Horizon
Memory
Open
Average
DM0.5 [ 31 ]
23.49 / 18.00
8.06 / 4.00
24.82 / 16.75
33.70 / 19.50
47.74 / 47.44
2.43 / 2.08
24.90 / 19.34
GalaxeaVLA (G0.5) [ 69 ]
26.74 / 20.00
11.16 / 6.00
28.25 / 20.42
44.12 / 32.25
8.61 / 7.33
1.73 / 1.58
20.23 / 14.88
Xiaomi-Robotics-1 [ 96 ]
35.65 / 28.00
11.44 / 6.00
26.69 / 18.83
38.39 / 23.67
7.81 / 6.56
3.94 / 3.58
20.07 / 13.93
Hy-Embodied-0.5-VLA [ 112 ]
21.98 / 17.00
1.57 / 0.00
13.81 / 8.00
25.74 / 14.92
13.37 / 12.11
0.65 / 0.58
13.07 / 8.80
Spatial Forcing [ 58 ]
21.25 / 15.00
6.98 / 4.00
17.33 / 10.58
23.26 / 14.58
5.43 / 4.11
1.78 / 1.58
12.38 / 8.04
π0.5 [ 45 ]
20.93 / 15.00
5.82 / 1.00
12.40 / 5.50
23.54 / 14.67
5.78 / 4.56
1.98 / 1.67
11.41 / 6.91
Table 3 : RoboDojo simulation results. Top 20 policies from the official August 24, 2026 leaderboard, ordered by average score. Each cell reports partial-progress score / success rate (%). Gen.-Std. and Gen.-Rand. denote the standard and randomized Generalization settings. The full leaderboard contains 35 policies; VLAct ranks eighth by score and sixth by success rate. The dark-blue row highlights VLAct; light-blue rows and † identify explicitly designated WAM entries.
Appendix figures & tables9 assets
Supplementary material from the paper’s appendix.
Appendix
Method
Safety
Distractor
Extrap.
Long-H.
Avg.
SmolVLA [ 78 ]
16.5
21.0
10.0
24.7
16.3
π0 -FAST [ 76 ]
34.1
39.0
4.2
20.7
25.6
GR00T-N1.6 [ 11 ]
32.8
40.7
16.7
10.3
27.8
Qwen3-VL- π [ 3 ]
38.3
45.3
22.4
25.3
34.1
UniVLA [ 90 ]
45.5
41.3
31.1
22.0
38.7
OpenVLA [ 54 ]
41.9
43.0
36.7
26.7
39.3
Appendix
Table 4: VLA-Arena results. Success rate (%) averaged over difficulty levels L0 / L1 / L2 within each task category. The final average follows the official suite weighting over 11 task suites. Dark blue highlights VLAct and per-column best results; light blue highlights the strongest public baseline by average and per-column second-best results.
Model
Backbone
SR ↑
MS ↑
OpenVLA
Llama-2
1.54
6.10
RDT-1B
DiT-1B
5.34
17.71
π0
PaliGemma
8.17
23.96
π0 -FAST
PaliGemma
3.54
20.87
π0.5
PaliGemma
9.63
26.17
InternVLA-M1
InternVL
5.40
27.57
Appendix
Table 5: DOMINO dynamic manipulation results. We evaluate one policy on all 35 DOMINO dynamic manipulation tasks under the clean dynamic setting. DOMINO reports Success Rate (SR) and Manipulation Score (MS). Dark blue highlights VLAct-OFT; light blue highlights the strongest baseline.
Pre-training update strategy
LIBERO-Plus
RoboTwin2.0
Update full backbone
78.9
77.1
Freeze vision encoder only
81.3
79.3
Freeze vision encoder + lower 1/2 LLM
82.6
80.5
Appendix
Table 6: Ablation on shallow-layer protection. Freezing the vision encoder and lower LLM layers during VLA pre-training reduces representation drift and improves downstream fine-tuning. Dark blue highlights the full design; light blue highlights the strongest partial variant.
Type
Dataset
Image Caption
LLaVA-ReCap-CC3M [ 56 ]
LLaVA OneVision [ 57 ]
BBox-QA
RefCOCO [ 21 ]
COCO-ReM [ 80 ]
Point-QA
PixMo-Points [ 30 ]
RoboPoint [ 110 ]
Appendix
Table 7: Auxiliary data used during VLA pre-training. The mixture covers captioning, grounding, spatial reasoning, and pure language instruction supervision.
Pre-training heads
Seen PI in pre-training
PI FT
Δ vs. scratch
None
–
60.5
–
OFT
No
55.1
− 5.4
OFT + GR00T
No
63.1
+ 2.6
OFT + PI + GR00T
Yes
77.0
+ 16.5
Appendix
Table 8: Head diversity reduces decoder lock-in. PI fine-tuning degrades after OFT-only pre-training, but improves after OFT+GR00T pre-training even though PI is excluded during pre-training. Light blue highlights the key unseen-head comparison; dark blue marks the full configuration with direct PI exposure.
Fine-tune head
From scratch
Single-head pre-train
Head-diverse pre-train
Δ vs. single
OFT
61.7
78.8
80.5
+ 1.7
PI
60.5
75.4
77.0
+ 1.6
GR00T
51.2
71.7
76.0
+ 4.3
Appendix
Table 9: Same-head adaptation under head-diverse pre-training. Head-diverse pre-training improves downstream performance even when the fine-tuning head matches one of the pre-training heads. Within each row, blue and lighter-blue cells mark the best and second-best results.
Setting
Unified Joint Space
Wrap Loss
Robotwin
Baseline
75.5
Unified joint space
✓
78.6
Unified joint space + Wrap-aware loss
✓
✓
80.5
Appendix
Table 10: Ablation of unified joint space and wrap-aware loss on RoboTwin. All models are trained under the base setting using 50 tasks with 50 trajectories per task, and evaluated on the Clean split. The pre-training and fine-tuning protocols are kept identical across all variants. Dark blue highlights the full design; light blue highlights the strongest partial variant.
Continued pre-training data
LIBERO-Plus
Original VLAct mixture
82.6
Original VLAct mixture + 20K RealOmin trajectories
83.7
Appendix
Table 11: Adding heterogeneous UMI trajectories to continued pre-training. Adding 20K trajectories from 10Kh-RealOmin-OpenData [ 37 ] , collected with a different embodiment and a UMI-style interface, improves downstream LIBERO-Plus performance. All other training and downstream fine-tuning settings are unchanged.
Setting
RoboTwin
LIBERO-Plus
Separate heads
78.5
81.1
Unified head
79.5
81.4
Unified action representation
80.5
82.6
Appendix
Table 12: Effect of unified action representation design. We compare separate heads, a unified head, and the full unified action representation on RoboTwin and LIBERO-Plus. Dark blue highlights the full design; light blue highlights the unified-head intermediate.