Scaling robot data is crucial for building generalist Vision-Language-Action (VLA) models, yet robot trajectories are harder to scale than web-scale image-text data because embodied collection is costly and sparsely covers the physical world. This makes representation quality a central bottleneck: under a fixed robot-data budget, continued pre-training must turn limited trajectories into transferable visual-action knowledge rather than merely fit actions. We propose VLAct, a VLA-oriented VLM backbone trained on broad, heterogeneous, multi-embodiment robot data before task-specific fine-tuning. VLAct preserves the broad VLM prior and encourages shared action semantics across embodiments through VLM-prior preservation, multi-head continuous action co-supervision, and a partially unified cross-embodiment action layout, while allowing task-specific action heads during fine-tuning. Across simulation, real-world, and unseen-embodiment transfer, VLAct consistently improves downstream performance under fixed fine-tuning protocols. On LIBERO-Plus and RoboTwin 2.0, VLAct surpasses industrial VLA systems including ABot-M0 and LingBot-VLA, achieving success rates of 82.6% and 92.5%. On RoboDojo, VLAct ranks sixth among all policies by success rate and outperforms all explicitly designated world-action model (WAM) entries on both metrics. Most notably, on RoboCasa-GR1, an unseen humanoid embodiment, VLAct using only 20% of downstream trajectories outperforms the full-data GR00T-N1.6 baseline. These results are obtained using fully open-source data and only a 16-GPU training setup, showing that representation-centric continued pre-training can deliver highly competitive performance under a modest compute budget and is an important independent axis of VLA progress beyond data scaling.
Figures & tables
Method
Camera
Robot
Lang.
Light
Bg.
Noise
Layout
Total
OpenVLA [ 54 ]
0.8
3.5
23.0
8.1
34.8
15.2
28.5
15.6
OpenVLA-OFT [ 53 ]
56.4
31.9
79.5
88.7
93.3
75.8
74.2
69.6
NORA [ 44 ]
2.2
37.0
65.1
45.7
58.6
12.8
62.1
39.0
WorldVLA [ 19 ]
0.1
27.9
41.6
43.7
17.1
10.9
38.0
25.0
UniVLA [ 90 ]
1.8
46.2
69.6
69.0
81.0
21.2
31.9
42.9
π0 [ 12 ]
13.8
6.0
58.8
85.0
81.4
79.0
68.8
53.6
Table 1 : Results on LIBERO-Plus. Per-dimension success rate (%) on the LIBERO-Plus robustness benchmark. All models are trained on standard LIBERO and evaluated on the perturbed test set. Dark blue highlights VLAct and per-column best results; light blue highlights the strongest prior method by total score and per-column second-best results.
Method
Family
Base Setting
Data Scaling Setting
Clean
Clean
Random
Published VLA and world-action-model systems
π0 [ 12 ]
Flow
46.4
65.9
58.4
π0.5 [ 45 , 95 ]
Flow
60.2
82.7
76.8
X-VLA [ 114 , 10 ]
Flow
70.0
72.8
72.8
Lingbot-VLA [ 95 ]
Flow
–
88.6
86.7
Table 2: RoboTwin 2.0 results. Task success rate (%) under the Base and Data Scaling settings. Base uses 50 clean trajectories per task and is evaluated on the Clean regime; Data Scaling additionally uses 500 randomized trajectories per task and is evaluated on both the Clean and Random regimes. Each reported regime evaluates 50 tasks with 100 episodes per task. Published numbers are taken from the cited papers or their same-protocol re-evaluations. Dark blue highlights our default VLAct-OFT configuration and light blue highlights alternative VLAct heads. Best results are bold and second-best results are underlined.
Method
Gen.-Std.
Gen.-Rand.
Precision
Long-Horizon
Memory
Open
Average
DM0.5 [ 31 ]
23.49 / 18.00
8.06 / 4.00
24.82 / 16.75
33.70 / 19.50
47.74 / 47.44
2.43 / 2.08
24.90 / 19.34
GalaxeaVLA (G0.5) [ 69 ]
26.74 / 20.00
11.16 / 6.00
28.25 / 20.42
44.12 / 32.25
8.61 / 7.33
1.73 / 1.58
20.23 / 14.88
Xiaomi-Robotics-1 [ 96 ]
35.65 / 28.00
11.44 / 6.00
26.69 / 18.83
38.39 / 23.67
7.81 / 6.56
3.94 / 3.58
20.07 / 13.93
Hy-Embodied-0.5-VLA [ 112 ]
21.98 / 17.00
1.57 / 0.00
13.81 / 8.00
25.74 / 14.92
13.37 / 12.11
0.65 / 0.58
13.07 / 8.80
Spatial Forcing [ 58 ]
21.25 / 15.00
6.98 / 4.00
17.33 / 10.58
23.26 / 14.58
5.43 / 4.11
1.78 / 1.58
12.38 / 8.04
π0.5 [ 45 ]
20.93 / 15.00
5.82 / 1.00
12.40 / 5.50
23.54 / 14.67
5.78 / 4.56
1.98 / 1.67
11.41 / 6.91
Table 3 : RoboDojo simulation results. Top 20 policies from the official August 24, 2026 leaderboard, ordered by average score. Each cell reports partial-progress score / success rate (%). Gen.-Std. and Gen.-Rand. denote the standard and randomized Generalization settings. The full leaderboard contains 35 policies; VLAct ranks eighth by score and sixth by success rate. The dark-blue row highlights VLAct; light-blue rows and † identify explicitly designated WAM entries.
Appendix figures & tables9 assets
Supplementary material from the paper’s appendix.
Appendix
Method
Safety
Distractor
Extrap.
Long-H.
Avg.
SmolVLA [ 78 ]
16.5
21.0
10.0
24.7
16.3
π0 -FAST [ 76 ]
34.1
39.0
4.2
20.7
25.6
GR00T-N1.6 [ 11 ]
32.8
40.7
16.7
10.3
27.8
Qwen3-VL- π [ 3 ]
38.3
45.3
22.4
25.3
34.1
UniVLA [ 90 ]
45.5
41.3
31.1
22.0
38.7
OpenVLA [ 54 ]
41.9
43.0
36.7
26.7
39.3
Appendix
Table 4: VLA-Arena results. Success rate (%) averaged over difficulty levels L0 / L1 / L2 within each task category. The final average follows the official suite weighting over 11 task suites. Dark blue highlights VLAct and per-column best results; light blue highlights the strongest public baseline by average and per-column second-best results.
Model
Backbone
SR ↑
MS ↑
OpenVLA
Llama-2
1.54
6.10
RDT-1B
DiT-1B
5.34
17.71
π0
PaliGemma
8.17
23.96
π0 -FAST
PaliGemma
3.54
20.87
π0.5
PaliGemma
9.63
26.17
InternVLA-M1
InternVL
5.40
27.57
Appendix
Table 5: DOMINO dynamic manipulation results. We evaluate one policy on all 35 DOMINO dynamic manipulation tasks under the clean dynamic setting. DOMINO reports Success Rate (SR) and Manipulation Score (MS). Dark blue highlights VLAct-OFT; light blue highlights the strongest baseline.
Pre-training update strategy
LIBERO-Plus
RoboTwin2.0
Update full backbone
78.9
77.1
Freeze vision encoder only
81.3
79.3
Freeze vision encoder + lower 1/2 LLM
82.6
80.5
Appendix
Table 6: Ablation on shallow-layer protection. Freezing the vision encoder and lower LLM layers during VLA pre-training reduces representation drift and improves downstream fine-tuning. Dark blue highlights the full design; light blue highlights the strongest partial variant.
Type
Dataset
Image Caption
LLaVA-ReCap-CC3M [ 56 ]
LLaVA OneVision [ 57 ]
BBox-QA
RefCOCO [ 21 ]
COCO-ReM [ 80 ]
Point-QA
PixMo-Points [ 30 ]
RoboPoint [ 110 ]
Appendix
Table 7: Auxiliary data used during VLA pre-training. The mixture covers captioning, grounding, spatial reasoning, and pure language instruction supervision.
Pre-training heads
Seen PI in pre-training
PI FT
Δ vs. scratch
None
–
60.5
–
OFT
No
55.1
− 5.4
OFT + GR00T
No
63.1
+ 2.6
OFT + PI + GR00T
Yes
77.0
+ 16.5
Appendix
Table 8: Head diversity reduces decoder lock-in. PI fine-tuning degrades after OFT-only pre-training, but improves after OFT+GR00T pre-training even though PI is excluded during pre-training. Light blue highlights the key unseen-head comparison; dark blue marks the full configuration with direct PI exposure.
Fine-tune head
From scratch
Single-head pre-train
Head-diverse pre-train
Δ vs. single
OFT
61.7
78.8
80.5
+ 1.7
PI
60.5
75.4
77.0
+ 1.6
GR00T
51.2
71.7
76.0
+ 4.3
Appendix
Table 9: Same-head adaptation under head-diverse pre-training. Head-diverse pre-training improves downstream performance even when the fine-tuning head matches one of the pre-training heads. Within each row, blue and lighter-blue cells mark the best and second-best results.
Setting
Unified Joint Space
Wrap Loss
Robotwin
Baseline
75.5
Unified joint space
✓
78.6
Unified joint space + Wrap-aware loss
✓
✓
80.5
Appendix
Table 10: Ablation of unified joint space and wrap-aware loss on RoboTwin. All models are trained under the base setting using 50 tasks with 50 trajectories per task, and evaluated on the Clean split. The pre-training and fine-tuning protocols are kept identical across all variants. Dark blue highlights the full design; light blue highlights the strongest partial variant.
Continued pre-training data
LIBERO-Plus
Original VLAct mixture
82.6
Original VLAct mixture + 20K RealOmin trajectories
83.7
Appendix
Table 11: Adding heterogeneous UMI trajectories to continued pre-training. Adding 20K trajectories from 10Kh-RealOmin-OpenData [ 37 ] , collected with a different embodiment and a UMI-style interface, improves downstream LIBERO-Plus performance. All other training and downstream fine-tuning settings are unchanged.
Setting
RoboTwin
LIBERO-Plus
Separate heads
78.5
81.1
Unified head
79.5
81.4
Unified action representation
80.5
82.6
Appendix
Table 12: Effect of unified action representation design. We compare separate heads, a unified head, and the full unified action representation on RoboTwin and LIBERO-Plus. Dark blue highlights the full design; light blue highlights the unified-head intermediate.
Vision-Language-Action (VLA) models widely adopt pretrained Vision-Language Models (VLMs) as policy backbones, yet it remains unclear what kind of pretrained VLM representation is useful as a VLA initialization. In this paper, we study VLA initialization as a controlled representation-design problem along three axes: capability-level embodied VQA supervision, parameter-update strategy, and robot-data pretraining. Our experiments show that the original pretrained VLM representation is a key source of action performance. However, embodied VQA adaptation does not yield uniform gains: its benefit depends on downstream bottlenecks, and gains from different capability domains are not simply additive. For update strategy, LoRA provides a more reliable initialization than Full Finetune, indicating that overly reshaping the pretrained representation can weaken VLA initialization. Robot-data pretraining further improves VLA initialization, with the strongest variant obtained by staged LoRA-based training. Together, these findings suggest that effective VLM-to-VLA adaptation should inject action-relevant embodied and robot-trajectory signals while preserving the pretrained VLM representation that remains useful for action learning.
Large-scale pretraining has made Vision-Language-Action (VLA) models promising foundations for generalist robot manipulation, yet adapting them to downstream tasks remains necessary. However, the common practice of full fine-tuning treats pretraining as initialization and can shift broad priors toward narrow training-distribution patterns. We propose PriorVLA, a novel framework that preserves pretrained priors and learns to leverage them for effective adaptation. PriorVLA keeps a frozen Prior Expert as a read-only prior source and trains an Adaptation Expert for downstream specialization. Expert Queries capture scene priors from the pretrained VLM and motor priors from the Prior Expert, integrating both into the Adaptation Expert to guide adaptation. Together, PriorVLA updates only 25% of the parameters updated by full fine-tuning. Across RoboTwin 2.0, LIBERO, and real-world tasks, PriorVLA achieves stronger overall performance than full fine-tuning and state-of-the-art VLA baselines, with the largest gains under out-of-distribution (OOD) and few-shot settings. PriorVLA improves over pi0.5 by 11 points on RoboTwin 2.0-Hard and achieves 99.1% average success on LIBERO. Across eight real-world tasks and two embodiments, PriorVLA reaches 81% in-distribution (ID) and 57% OOD success with standard data. With only 10 demonstrations per task, PriorVLA reaches 48% ID and 32% OOD success, surpassing pi0.5 by 24 and 22 points, respectively.
Xinyu Guo, Bin Xie, Wei Chai +4
Institute of Automation, Chinese Academy of Sciences · Dexmal · Nanjing University of Aeronautics and Astronautics +1
Large-scale Vision-Language-Action (VLA) pretraining is increasingly adopted as the foundation for robot policies, yet the evidence for pretrained VLAs is almost invariably reported after task-specific fine-tuning. This leaves a foundational question unanswered: does VLA pretraining itself yield executable robot behavior, or does it merely furnish a better initialization for downstream policy learning? We present Wall-OSS-0.5, an open-source 4B VLA built upon a 3B VLM backbone augmented with action-generation components, designed so that pretrained robotic capability is directly measurable on physical hardware. The model is pretrained across more than 20 embodiments, processing over one million robot trajectories per epoch alongside a grounded multimodal corpus. We adopt a gradient-bridged co-training recipe in which three objectives play distinct and complementary roles: discrete action prediction routes strong VLM-native gradients into the backbone, multimodal prediction preserves grounded vision-language understanding, and continuous flow matching serves as the deployment-time action interface. Before task-specific fine-tuning, the pretrained checkpoint achieves non-trivial zero-shot real-robot behavior, completing several tasks, including a held-out deformable manipulation task, at high task progress on a 17-task suite. After fine-tuning, the same checkpoint serves as a stronger adaptation prior, reaching 60.5% average task progress on 15 real-robot tasks and outperforming π_0.5 by 17.5%. Multimodal evaluations further confirm that action training does not erode grounded vision-language competence: the model preserves broad vision-language ability while strengthening embodied grounding. Together, these results reposition VLA pretraining from an initialization strategy to a directly testable, already useful source of robot capability.